Systems and programs, etc.
Patent Information
- Application Number
- JP2025032232
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0158】 本発明によれば、従来よりも優れたシステム等を提供できる。
Smart Images

Figure 2026144753000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to, for example, systems and programs. [Background Art]
[0002] An image processing system that displays captured images taken by an on-board camera mounted on a work vehicle is known in the art. The image processing system described in Patent Document 1 is a system that performs planarization processing on captured images and displays the processed images. [Prior Art Documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Unexamined Patent Application Publication No. 2017-132298 [Summary of the Invention] [Problems to be Solved by the Invention]
[0004] However, conventional systems have had various problems. Therefore, an object of the present invention is to provide a system and a program that have superior characteristics compared to conventional systems and programs.
[0005] The object of the present invention is not limited thereto, and the applicant intends to obtain rights through divisional applications, amendments, etc., for configurations that aim to obtain the effects derived from the components of the configuration disclosed in this specification and the drawings, etc. For example, problems that can be described in this specification as "~is possible" or "~is feasible" are disclosed in this specification. Each problem is described independently, and the applicant intends to obtain rights to each configuration for solving each problem independently through divisional applications, amendments, etc. Even if a problem is implicitly understood from the description in the specification, the applicant intends to include a part of the configuration described in this specification in the claims through amendment or divisional application. Furthermore, configurations that solve problems by combining these independent problems are also disclosed, and the applicant intends to obtain rights to them. [Means for solving the problem]
[0006] (1) For example, the image processing system may include an image processing means that extracts one or more image regions such that a portion of the input image overlaps, and converts the extracted one or more image regions into an output image with a different region shape from the image regions.
[0007] This configuration yields an output image in which one or more image regions extracted from the input image, with some overlapping areas, are transformed into different regional shapes. For example, the overlapping area can be set to ensure the visibility of the subject in the input image. In this way, the visibility of the subject in the input image can be ensured.
[0008] Furthermore, it is preferable to include a configuration that performs the process of outputting an output image. For example, a configuration that performs the process of outputting an output image may include a configuration that performs the process of displaying the output image. It is preferable to include a display means that is displayed by the display process. It is preferable to include a configuration that performs the process of displaying the management screen on the display means by the display process.
[0009] For example, the output image may be displayed on the user's management screen. It is preferable to have a configuration that processes the display so that subjects in the overlapping area are not cut off in the output image. In this way, subjects in the overlapping area are not cut off in the output image, allowing for smooth management work using the output image.
[0010] The input image could, for example, be a spherical image obtained from a spherical camera mounted on a work vehicle. In this way, output images with different regional shapes can be obtained from spherical images that capture a wide area around the work vehicle.
[0011] For example, the image region may be arc-shaped, and the region shape may be rectangular. In this way, an output image can be obtained in which the arc-shaped image region obtained from the celestial sphere image has been transformed into a rectangular region shape. In this way, an output image can be obtained in which the arc-shaped image region, which is part of the celestial sphere image, has been unfolded into a rectangular shape, and by viewing the output image, the visibility of the subject captured in the celestial sphere image can be improved.
[0012] For example, the system may include a configuration that sets boundaries for areas that are blind spots for the on-board camera, such as the pillars of a work vehicle. In this way, the image region defined by the blind spot of the on-board camera is extracted from the celestial image, and the image region can be effectively defined.
[0013] For example, the system may include a configuration for extracting a single image region. Alternatively, it may include a configuration for setting overlapping areas of different sizes for multiple image regions. For example, the system may include a configuration for extracting three or more image regions. In this way, the output image can be subdivided and the region shape transformed, improving the visibility of the subject captured in the entire output image.
[0014] Furthermore, it is preferable to have a configuration in which multiple image regions are of different sizes. Also, it is preferable to have a configuration in which the region shape, distinct from the image regions, is, for example, rectangular. This ensures the visibility of the output image.
[0015] Furthermore, for the conversion to the output image, it is preferable to have a configuration that performs a process such as performing an equirectangular transformation on a circular celestial sphere image. In this way, a rectangular equirectangular image can be obtained from a circular celestial sphere image.
[0016] Furthermore, the image processing means may include, for example, a configuration that displays the output image on a display unit in real time. In this way, the user can check the output image in real time.
[0017] Furthermore, the image processing system should ideally include a configuration that outputs the output image to multiple terminals. This would allow the output image to be viewed on multiple terminals, improving convenience. (2) The image processing means may be configured to generate a display image by arranging a plurality of the output images. In this way, a display image consisting of multiple output images arranged side by side can be obtained.
[0018] Furthermore, for example, it is desirable to have a configuration that processes the display of the display image on the display unit. In this way, a display image consisting of multiple output images arranged in a row can be displayed on the display unit.
[0019] For example, the image processing means may be configured to generate a display image in which the two output images are displayed side by side, one above the other. In this way, the user can view the two output images displayed side by side in one display image at once, improving the user's ability to view both output images at a glance.
[0020] (3) The input image may comprise an image captured in spherical coordinates and projected onto polar coordinates, and the image processing means may preferably be configured to extract the image region from an annular region of the input image excluding a circular region at the coordinate center.
[0021] According to this configuration, an image region is extracted from an annular region of the input image excluding the central circular region. For example, the input image may preferably comprise a celestial image in which the information density at the central portion tends to be low. According to this configuration, an image region can be extracted from the annular region excluding the central portion with low information density in the celestial image, and the information density of the extracted image region can be increased.
[0022] Furthermore, since the image extraction means performs processing excluding the central circular region in the input image, it is possible to reduce the processing load and improve the processing speed in subsequent region shape conversion.
[0023] (4) The image processing means may preferably be configured to extract the image region by adjusting the ratio of a radial length in the annular region to a length of an arc located outward in the radial direction, based on a preset aspect ratio of the display unit for the display image.
[0024] According to this configuration, the length ratio in the annular region is adjusted based on the aspect ratio of the display unit, and the image region is extracted. For example, when a display unit having a known aspect ratio is used, the input image may preferably comprise images of a plurality of sizes.
[0025] According to this configuration, regardless of the size of the input image, an image region having a size that matches the preset aspect ratio of the display unit can be extracted. That is, it is possible to eliminate the need for the user to individually adjust the length ratio in the annular region for each input image of different sizes in accordance with the size of the display unit, thereby improving user convenience. (5) The image processing means may preferably be configured to accept setting of a range of the overlapping portion in the image region based on a specification from a user.
[0026] With this arrangement, the range of the overlapping portion in the image area is set based on the specification received from the user. For example, a configuration including an input unit for a user to perform an input operation for specifying the range of the overlapping portion is preferable.
[0027] With this arrangement, when the user performs an input operation on the input unit, the user can arbitrarily set the range of the overlapping portion according to the application and purpose, and the convenience for the user can be ensured. Further, for example, a configuration that performs processing for displaying the overlapping range specified by the user on a display unit is preferable. In this case, a configuration that performs processing for displaying the overlapping range specified by the user so as to be superimposed on the input image displayed on the display unit is preferable. With this arrangement, the user can accurately grasp the range of the overlapping portion that the user himself / herself specifies for the input image. (6) It is preferable that the image processing means comprises a configuration for changing the range of the overlapping portion based on a preset attribute of a detection target object.
[0028] With this arrangement, the range of the overlapping portion is changed based on a preset attribute of the detection target object. For example, a configuration including an input unit for a user to perform an input operation for setting the attribute of the detection target object in advance is preferable.
[0029] With this arrangement, when the user operates the input unit to set the detection target object in advance, the range of the overlapping portion where the entire detection target object is sufficiently captured can be adaptively changed. This can prevent the problem that when the detection target object is large, the detection target object is not sufficiently captured due to the small range of the overlapping portion. (7) It is preferable that the image processing means comprises a configuration for changing the range of the overlapping portion based on an attribute of a subject captured in the input image.
[0030] In this way, the extent of the overlapping area is changed based on the attributes of the subject captured in the input image. For example, it is preferable to have a configuration that performs processing to determine the attributes of the subject captured in the input image. In this way, the extent of the overlapping area that sufficiently captures the entire subject can be adaptively changed in real time according to the subject captured in the input image, resulting in a highly responsive system.
[0031] Furthermore, the scope of the overlapping area should be configured so that the user can pre-set it for each attribute of the subject. In this way, the user can set the scope of the overlapping area for each attribute of the subject, thereby increasing its suitability for various applications. (8) The image processing means may be configured to change the range of the overlapping portion based on the position of the subject captured in the input image.
[0032] In this way, the extent of the overlapping area is changed based on the position of the subject in the input image. For example, it is preferable to have a configuration that performs a process to set a predetermined overlapping area when the subject is located near the boundary of the default image area. In this way, the extent of the overlapping area can be adaptively changed in real time based on the position of the subject in the input image, for example, when the subject is located near the boundary of the image area, and the extent of the overlapping area can be appropriately changed in consideration of the movement of the subject. (9) The image processing means may be configured to change the range of the overlapping portion based on the shooting range of the input image.
[0033] In this way, the extent of the overlapping area is changed based on the capture range of the input image. For example, it is preferable to have a configuration that processes the capture range of the input image from the input image. In this way, the extent of the overlapping area of the image region can be adaptively changed based on the detected capture range of the input image, and the versatility of the image processing system can be increased regardless of the specifications of the camera that captures the input image.
[0034] Furthermore, it is preferable to include a configuration that allows the user to pre-set the shooting range of the input image. In this way, the range of overlapping areas can be changed to an appropriate value based on the shooting range of the input image pre-set by the user according to the application of the image processing system.
[0035] (10) The image processing means further comprises a detection means for detecting a pre-set object from the input image, and when the detection means detects the object, it is preferable that the image processing means is configured to output information indicating the object in the area of the object in the display image.
[0036] In this way, when the detection means detects a pre-set target object from the input image, information indicating the detected object is output to the area of the detected object in the displayed image. For example, the system may include a configuration that displays an emphasis frame that highlights the area in which the detected object is captured, as information indicating the detected object.
[0037] Furthermore, it is preferable to include a configuration that processes text data indicating the attributes of the detected object (person, vehicle, material, etc.) as information indicating the detected object. In this way, information indicating the detected object (highlighted frame, object attributes, etc.) is output in the area of the displayed image in which the detected object detected from the input image is captured, allowing the user to immediately grasp the detection result from the detection means along with its location.
[0038] Furthermore, it is preferable to have a configuration that, for example, when the same object is detected in the overlapping portion of multiple output images within the displayed image, outputs information indicating the detected object in one output image and in each of the other output images. In this way, the user who has viewed the displayed image can reliably understand the detection result of the object.
[0039] (11) The system further comprises a detection means for detecting a pre-set object from the input image, wherein the input image is captured by an on-board camera mounted on the work vehicle, and the detection means is configured to detect the presence of the driver in the work vehicle by reading a predetermined worker code assigned to the driver of the work vehicle.
[0040] In this way, the detection means reads a predetermined worker code assigned to the driver from the input image captured by the on-board camera mounted on the work vehicle, thereby detecting the presence of the driver in the work vehicle as a pre-set detection target.
[0041] For example, it would be beneficial to have a configuration in which worker codes and driver information are linked and stored together. In this way, the detection means can read the worker code assigned to the driver from the input image, which is from the on-board camera of the work vehicle, and detect the presence of the driver in the work vehicle, thereby accurately detecting the information of the driver who is actually operating the work vehicle.
[0042] Furthermore, it is preferable to have a configuration in which the identification information of a known work vehicle and the worker code of the driver assigned to that work vehicle are linked and stored. It is also preferable to have a configuration in which the detection means performs a process of matching the worker code detected with the identification information of the work vehicle. In this way, when the presence of a driver is detected from the input image of the work vehicle's onboard camera, it is possible to determine whether the appropriate worker is driving that work vehicle.
[0043] (12) The detection means may be configured to change the evaluation region in the input image for detecting the presence of the driver based on the position of the on-board camera in the work vehicle.
[0044] In this way, the detection means changes the evaluation area for detecting the presence of the driver based on the position of the onboard camera in the work vehicle. For example, if the work vehicle is of a type in which the onboard camera is mounted at the top center of the interior, it is preferable to have a configuration that performs a process to set the evaluation area in the input image to the center, assuming that the driver is in the center of the input image.
[0045] Furthermore, for example, if the work vehicle is of a type in which an on-board camera is mounted on the upper center of the interior, it would be beneficial to have a configuration that processes the input image to set the evaluation area to be further back in the vehicle than the center, assuming that the driver is located towards the rear of the vehicle. In this way, even if the mounting position of the on-board camera differs due to differences in the type of work vehicle, for example, the appropriate evaluation area can be automatically set, ensuring user convenience.
[0046] (13) The detection means may be configured to change the evaluation region in the input image for detecting the presence of the driver based on the positional relationship with a specific structure fixed to and used on the work vehicle.
[0047] In this way, the detection means changes the evaluation area for detecting the presence of the driver based on the positional relationship with a specific structure that is fixed to and used on the work vehicle. For example, the system may be configured to include a process for setting the specific structure as a structure that obstructs the view from the driver's seat of the work vehicle, such as a mast or head guard when the work vehicle is a forklift.
[0048] In this way, areas that become blind spots due to the presence of specific structures within the field of view of the on-board camera mounted on the work vehicle can be excluded from the evaluation area, thereby reducing the load on the detection process.
[0049] (14) The detection means may be configured to perform detection processing on each of the multiple input images obtained from each of the multiple in-vehicle cameras and to detect the presence of the driver based on the respective detection results.
[0050] In this way, the detection means performs detection processing on multiple input images obtained from multiple in-vehicle cameras, and then combines the detection results from both to detect the presence of the driver. For example, it is preferable that the multiple in-vehicle cameras are set up so that each has a different shooting range. In this way, the presence of the driver can be detected simultaneously from a wide shooting range obtained from multiple in-vehicle cameras.
[0051] (15) The detection means may be configured to output the detection results when the detection result for the input image from the first in-vehicle camera and the detection result for the input image from the second in-vehicle camera are within a predetermined range.
[0052] In this way, the detection means outputs the detection result when the detection result from the first in-vehicle camera and the detection result from the second in-vehicle camera match within a predetermined range. For example, it is preferable to have a configuration that outputs the detection results when the position in the input image where the object to be detected is within a predetermined range in each detection result. In this way, the accuracy of the estimated position of the detected object can be improved compared to when an input image obtained from a single camera is used.
[0053] (16) The detection means may be configured to detect the presence of the driver when the relative position of the driver with respect to the surrounding structure detected with respect to the input image from the first in-vehicle camera and the relative position of the driver with respect to the surrounding structure detected with respect to the input image from the second in-vehicle camera are within a predetermined range.
[0054] In this way, the detection means detects the presence of a driver if the driver's relative position to surrounding structures is within a predetermined range based on the output images obtained from the first and second on-board cameras, respectively.
[0055] For example, it is preferable to have a configuration that processes the setting of surrounding structures such as a head guard when the work vehicle is a forklift, where the relative position to the driver does not change easily. In this way, it is possible to perform highly accurate detection that takes into account the relative positional relationship with surrounding structures using the output images obtained from the first and second on-board cameras, respectively.
[0056] (17) The image processing system may further include a notification means, wherein the notification means may be configured to notify the detection result of the object to be detected when the presence of the driver in the work vehicle is detected by the detection means.
[0057] In this way, the notification means will notify the detection result of the object when a driver is present in the work vehicle. For example, the notification means may be configured to notify the detection result by outputting an alert sound. Alternatively, the notification means may be configured to notify the detection result by outputting an alert display on the display unit. Furthermore, the notification means may be configured to notify the detection result by flashing a warning light. In this way, the driver can be reliably informed of the detection results of objects around the work vehicle.
[0058] (18) The notification means may be configured to notify the detection result of the object to be detected in a different manner than when the presence of the driver is detected by the detection means, when the presence of the driver is not detected by the detection means.
[0059] In this way, the notification means will notify the detection result of the object to be detected in different ways depending on whether or not a driver is present in the work vehicle. For example, the system may be configured to output an alert to the display unit inside the vehicle when a driver is present in the work vehicle, and to flash an external warning light when a driver is not present in the work vehicle. In this way, by checking the notification method of the detection result, workers positioned around the work vehicle can know whether or not there is a driver in the work vehicle.
[0060] (19) The image processing system further comprises notification means, and the notification means may be configured to provide different ways of notifying the detection result of the detected object based on the positional relationship of the detected object with respect to the work vehicle when the detection means detects the detected object in the vicinity of the work vehicle.
[0061] In this way, the notification means will vary the manner in which it notifies the detection result of the object based on the positional relationship of the object to the work vehicle. For example, it is preferable to have a configuration that sets up a stronger alert when the work vehicle and the object are close together compared to when the work vehicle and the object are far apart.
[0062] In this way, if the object to be detected is found in close proximity to the work vehicle, a stronger alert can be issued than if it is detected at a distance from the work vehicle, effectively drawing attention to the surrounding area.
[0063] (20) The notification means may be configured to notify the detection result of the object in the first manner if the positional relationship between the object to be detected and the work vehicle is within a predetermined range that has been arbitrarily set in advance, and to leave the detection result of the object to be detected unattended in the second manner if it is outside the predetermined range.
[0064] In this manner, the notification means notifies the detection result in the first mode if the positional relationship between the object to be detected and the work vehicle is within a predetermined range, and notifies the detection result in the second mode if it is outside the predetermined range.
[0065] For example, in the first embodiment, an alert sound and alert display may be output to the display unit inside the vehicle, and in the second embodiment, a configuration may be provided that causes a warning light outside the vehicle to flash. In this way, the notification method of the detection result of the detected object can be made different depending on the positional relationship between the work vehicle and the detected object. For example, when the work vehicle and the detected object are close together, a stronger alert can be issued to draw attention to the surroundings compared to when the work vehicle and the detected object are far apart.
[0066] (21) The image processing system further comprises notification means, and the notification means may further comprise a configuration that causes the notification of the detection result of the object to be detected to vary based on at least one of the status of the work vehicle, the presence or absence of an object to be detected in the vicinity of the work vehicle, and the location of the object to be detected.
[0067] In this way, the notification means will vary the notification method of the detection result of the object based on at least one of the status of the work vehicle, the presence or absence of the object to be detected in the vicinity of the work vehicle, and the location of the object to be detected.
[0068] For example, the system may include a configuration that uses the operating information of the work vehicle (such as CAN information) as the status of the work vehicle. Furthermore, it may include a configuration that sets the location of the object to be detected in the area behind the work vehicle, which is particularly dangerous. In this way, the notification method for the detection result of the object can be changed based on various environmental factors surrounding the work vehicle, enabling a flexible notification function tailored to the use of the work vehicle.
[0069] (22) The computer may include a method for performing an image processing step in which it extracts one or more image regions such that some areas of the input image overlap, and converts the extracted one or more image regions into an output image with a different region shape from the image regions. In this way, each step of the image processing system described in (1) can be performed. (23) The computer may be a program that causes it to function as one of the means in the image processing system described in any of (1) to (21).
[0070] In this way, by having the processor execute the image processing program, the functions of each means of the image processing system described in any one of claims 1 to 21 can be realized.
[0071] (1-1) The image processing system comprises an image extraction means for extracting one or more image regions from a captured celestial spherical image, and an image conversion means for converting the extracted image regions into a planar image, wherein the image extraction means extracts the image regions such that at least a portion of the celestial spherical image overlaps with the planar image. In this way, image regions can be extracted such that at least a portion of the celestial spherical image overlaps, and then converted into a planar image.
[0072] It is advisable to configure the system to use the extracted data for image processing. In this way, by overlapping a portion of the image region extracted from the celestial image, the visibility after conversion to a planar image can be improved. It is also advisable to configure the system to use the extracted data for display. In this way, by displaying a planar image that includes the overlapping image region, smooth visual recognition becomes possible.
[0073] It is advisable to configure the system so that the extracted data is used for AI processing. In this way, the AI can analyze planar images that include overlapping image regions, thereby improving the accuracy of feature extraction and object recognition.
[0074] (1-2) The image extraction means may further include an image synthesis unit that extracts multiple image regions, the image conversion means may convert multiple planar images, and the image processing system may further include an image synthesis unit that generates a display image by arranging the multiple planar images. In this way, multiple planar images can be converted and a display image by arranging them can be generated. The generated image may be used for image processing.
[0075] In this way, by generating a display image that arranges multiple planar images, a wide area of the image can be efficiently viewed. It is advisable to configure the system to use the generated image for display. In this way, by displaying multiple planar images side by side, information from a wide area can be grasped at once. It is advisable to configure the system to use the generated image for AI processing. In this way, by having the AI comprehensively analyze multiple planar images, a wide area of image information can be analyzed in a unified manner.
[0076] (1-3) The image extraction means may extract multiple arc-shaped portions from the annular region of the celestial sphere image, excluding the circular region of the coordinate center, as the image regions. In this way, multiple arc-shaped portions can be extracted from the annular region excluding the circular region of the coordinate center. The system may be configured to use the multiple extracted portions for image processing.
[0077] In this way, by extracting arc-shaped portions from a circular region, images of specific directions or ranges can be effectively utilized. It is advisable to use a configuration that displays multiple extracted portions. In this way, by displaying arc-shaped images extracted from a circular region, specific ranges can be emphasized. It is advisable to use a configuration that uses multiple extracted portions for AI processing. In this way, by extracting arc-shaped images, the accuracy of AI analysis regarding specific directions can be improved.
[0078] (1-4) The image extraction means may set the diameter of the circular region based on the aspect ratio of the display portion of the displayed image, and adjust the ratio of the diameter of the arc-shaped portion to the length of the arc. In this way, the diameter of the circular region can be set and the ratio of the arc-shaped portion can be adjusted based on the aspect ratio of the display portion of the displayed image. The adjusted image may be used for image processing. In this way, by adjusting the size of the circular area according to the aspect ratio of the display, an appropriate screen configuration can be achieved. It is best to use the adjusted configuration for display.
[0079] In this way, by adjusting the circular area while considering the aspect ratio, the image can be displayed in the optimal layout. It is advisable to use the adjusted image for AI processing. In this way, by adjusting the circular area while considering the aspect ratio, the uniformity of the training data for AI can be improved.
[0080] (1-5) The image extraction means may set the range of the overlapping portion in the image region according to the user's specifications. In this way, the range of the overlapping portion of the image region can be set according to the user's specifications. The set range should be used for image processing.
[0081] This approach allows users to freely change the extent of overlapping areas in the image region, improving customization. It is advisable to configure the system to use the configured settings for display. This allows the system to adjust the extent of overlapping areas based on user settings, appropriately displaying necessary information. It is also advisable to configure the system to use the configured settings for AI processing. This allows the AI to perform adaptive learning more easily by adjusting the overlapping areas based on user settings.
[0082] (1-6) The image extraction means may set the range of the overlapping portion based on the attributes of the object to be detected, which have been set in advance. In this way, the range of the overlapping portion of the image area can be set based on the attributes of the object to be detected. The set range should be used for image processing.
[0083] In this way, by setting the overlapping range according to the attributes of the detected object, the appropriate image region can be extracted. It is advisable to configure the system to use the configured region for display. In this way, by adjusting the overlapping range according to the attributes of the detected object, a highly visible image can be provided. It is advisable to configure the system to use the configured region for AI processing. In this way, by adjusting the overlapping range according to the attributes of the detected object, the AI's identification accuracy can be improved.
[0084] (1-7) The image extraction means may set the range of the overlapping portion based on the attributes of subjects captured near the boundaries of the pre-set image regions. In this way, the range of the overlapping portion of the image regions can be set based on the attributes of subjects captured near the boundaries of the image regions. The set range should be used for image processing.
[0085] In this way, by setting the overlapping range according to the attributes of the subject near the boundary of the image area, information from the necessary parts can be secured. It is best to configure the system to use the configured settings for display. In this way, by considering the attributes of the subject that appears near the boundary of the image area, information can be displayed more clearly. It is best to configure the system to use the configured settings for AI processing. In this way, by considering the attributes of the subject near the boundary of the image area, the AI can prioritize learning important subjects.
[0086] (1-8) The image extraction means may set the range of the overlapping portion based on the position of the subject captured in the celestial spherical image. In this way, the range of the overlapping portion of the image area can be set based on the position of the subject captured in the celestial spherical image. The set range should be used for image processing. In this way, important objects can be appropriately captured by setting the overlapping range considering the position of the subject captured in the celestial spherical image. The set range should be used for display.
[0087] In this way, by adjusting the overlapping area while considering the position of subjects in the celestial image, it is possible to generate images with high visibility. It is recommended to configure the system so that the settings are used for AI processing. In this way, the spatial recognition capabilities of the AI can be enhanced by adjusting the overlapping area while considering the position of subjects in the celestial image.
[0088] (1-9) The image extraction means may set the range of the overlapping portion based on the shooting range of the celestial spherical image. In this way, the range of the overlapping portion of the image area can be set based on the shooting range of the celestial spherical image. The set range should be used for image processing.
[0089] This approach enables optimal image extraction by setting overlapping areas based on the shooting range. It is advisable to configure the system to use the configured settings for display. This way, appropriate information is displayed on the screen through image adjustment based on the shooting range. It is also advisable to configure the system to use the configured settings for AI processing. This way, the AI's analysis target is appropriately limited through image adjustment based on the shooting range, eliminating unnecessary data.
[0090] (1-10) The system further includes a display unit that displays the generated display images, and it is preferable that the display unit be able to display multiple display images. In this way, multiple display images can be displayed on the display unit. It is preferable that the displayed images be used for image processing.
[0091] In this way, by displaying multiple images on the display unit, information from different perspectives can be viewed simultaneously. It is advisable to configure the system so that the displayed images are used for display purposes. In this way, by displaying multiple images simultaneously, information from different perspectives can be easily compared. It is advisable to configure the system so that the displayed images are used for AI processing. In this way, by having the AI integrate and analyze multiple displayed images, highly accurate analysis combining information from multiple perspectives becomes possible.
[0092] (1-11) The display unit further includes a detection means for detecting a pre-set object from the displayed image, and when the detection means detects the object, the display unit may superimpose a rectangular frame over the area of the object in the displayed image. In this way, the object can be detected from the displayed image and the rectangular frame can be superimposed. The superimposed image may be used for image processing.
[0093] In this way, the rectangular frame display of the detected object allows for an intuitive understanding of its location. It is advisable to use a configuration that uses the superimposed display for visualization. In this way, highlighting the detected object with a rectangular frame enables visually intuitive recognition. It is advisable to use a configuration that uses the superimposed display for AI processing. In this way, clearly indicating the detected object with a rectangular frame improves the accuracy of data annotation by AI.
[0094] (1-12) The system further comprises a detection means for detecting a pre-set object from the displayed image, wherein the spherical image is captured by an on-board camera mounted on the work vehicle, and the detection means detects the presence of the worker in the work vehicle by reading a predetermined worker code assigned to the driver of the work vehicle.
[0095] In this way, a spherical image can be captured by an in-vehicle camera, and the presence of workers can be detected by reading the worker code. It is advisable to configure the system to use the detected data for image processing. In this way, the presence of a specific worker can be accurately detected by reading the worker code. It is advisable to configure the system to use the detected data for display. In this way, it is possible to read the worker code and display the presence of the worker. It is advisable to configure the system to use the detected data for AI processing. In this way, by having the AI recognize the worker code, automatic individual identification and behavioral analysis become possible.
[0096] (1-13) The detection means may change the evaluation area in the displayed image that detects the presence of the worker based on the position of the on-board camera in the work vehicle. In this way, the evaluation area that detects the presence of the worker can be changed based on the position of the on-board camera. The modified area may be used for image processing.
[0097] In this way, by changing the evaluation area according to the position of the in-vehicle camera, detection suitable for the shooting environment becomes possible. It is advisable to configure the system to use the modified area for display. In this way, by changing the evaluation area according to the camera position, information at the appropriate location can be highlighted. It is advisable to configure the system to use the modified area for AI processing. In this way, by changing the evaluation area according to the camera position, the AI can learn from the optimal viewpoint.
[0098] (1-14) The detection means may change the evaluation area in the displayed image that detects the presence of the worker based on the positional relationship with a specific structure fixed to the work vehicle. In this way, the evaluation area that detects the presence of the worker can be changed based on the positional relationship with a specific structure. The modified area may be used for image processing.
[0099] In this way, by changing the evaluation area based on its positional relationship with structures, optimal detection for the target environment can be achieved. It is advisable to configure the system to use the modified area for display. In this way, the display of the target area can be optimized according to its positional relationship with specific structures. It is advisable to configure the system to use the modified area for AI processing. In this way, by considering the positional relationship with specific structures, the AI can perform contextual analysis more easily.
[0100] (1-15) The detection means may individually perform detection processing on each of the multiple display images obtained from each of the multiple in-vehicle cameras and detect the presence of the worker based on the respective detection results. In this way, detection processing can be performed on each of the display images from multiple in-vehicle cameras and the presence of the worker can be detected based on the detection results. It is preferable to configure the system so that the detected images are used in image processing.
[0101] In this way, by integrating detection results from multiple cameras, highly accurate worker detection becomes possible. It is advisable to configure the system to use the detected data for display. In this way, results from multiple cameras are integrated, and accurate worker detection information can be displayed. It is advisable to configure the system to use the detected data for AI processing. In this way, the AI integrates the results from multiple cameras, enabling improved accuracy through data fusion from different perspectives.
[0102] (1-16) The detection means may adopt the detection result if the detection result for the display image from the first in-vehicle camera and the detection result for the display image from the second in-vehicle camera are within a predetermined range. In this way, if the detection result of the first in-vehicle camera and the detection result of the second in-vehicle camera match within a predetermined range, the detection result can be adopted. The adopted result may be used for image processing.
[0103] This approach reduces false positives by comparing detection results from different cameras and only using those that match. It is advisable to configure the system to use the selected results for display. This reduces misrecognition by only displaying matching detection results from different cameras. It is advisable to configure the system to use the selected results for AI processing. This reduces AI misrecognition by comparing detection results from different cameras and only using those that match.
[0104] (1-17) The detection means may detect the presence of a worker when the relative position of the worker with respect to the surrounding structure detected in the display image from the first in-vehicle camera and the relative position of the worker with respect to the surrounding structure detected in the display image from the second in-vehicle camera show a degree of agreement within a predetermined range.
[0105] In this way, based on the display images from the first and second on-board cameras, the presence of a worker can be detected if their relative position to surrounding structures matches. It is preferable to configure the system to use the detected information for image processing. By considering the relative position of each camera during detection, more reliable results can be obtained.
[0106] It is advisable to configure the system to use the detected objects for display. This allows for the display of highly reliable detection results, taking into account the relative positions of each camera. Alternatively, it is advisable to configure the system to use the detected objects for AI processing. This allows the AI to appropriately understand the relationships between viewpoints by considering the relative positions of each camera.
[0107] (1-18) The system further includes a notification unit that notifies the detection result of the object by the detection means, and the notification unit notifies the detection result of the object when the presence of the worker in the work vehicle is detected by the detection means. In this way, the detection result of the object can be notified when the presence of the worker is detected in the work vehicle. The notified information may be used for image processing.
[0108] In this way, by notifying the detection result of an object while an operator is present, appropriate warnings can be provided. It is advisable to configure the system to use the notified information for display. In this way, the detection result when an operator is present can be clearly displayed, prompting appropriate warnings. It is advisable to configure the system to use the notified information for AI processing. In this way, the AI can detect an object while an operator is present and automate appropriate actions.
[0109] (1-19) The notification unit may notify the detection result of the object in a different manner than when the worker's presence is detected, when the worker's presence is not detected by the detection means. In this way, when the worker's presence is not detected in the work vehicle, the detection result of the object can be notified in a different manner. The notified information may be used for image processing.
[0110] In this way, by changing the notification method when no worker is present, situation-appropriate warnings can be implemented. It is advisable to configure the system to use the notified information for display. In this way, when no worker is present, the method of displaying the warning can be changed to ensure appropriate attention. It is advisable to configure the system to use the notified information for AI processing. In this way, the AI can issue different warnings when no worker is present, enabling a situation-appropriate response.
[0111] (1-20) The notification unit may, when an object is detected around the work vehicle by the detection means, provide different notifications of the object detection result depending on the positional relationship between the object and the work vehicle. In this way, when an object is detected around the work vehicle, the notification method can be changed depending on the positional relationship between the object and the work vehicle. The notification should be used for image processing.
[0112] In this way, the danger can be appropriately conveyed by changing the notification method according to the positional relationship between the object and the work vehicle. It is preferable to configure the system to use the notified information for display. In this way, appropriate information can be presented by changing the notification based on the positional relationship between the object and the work vehicle. It is preferable to configure the system to use the notified information for AI processing. In this way, an appropriate warning system can be constructed by having the AI adjust the notification method according to the positional relationship between the object and the work vehicle.
[0113] (1-21) The notification unit may notify the detection result of the object in the first manner if the positional relationship between the object and the work vehicle is within a predetermined range arbitrarily set in advance, and notify the detection result of the object in the second manner if it is outside the predetermined range. In this way, if the positional relationship between the object and the work vehicle is within the predetermined range, notification can be given in the first manner, and if it is outside the range, notification can be given in the second manner.
[0114] It is advisable to configure the system to use the notified information for image processing. In this way, safety can be ensured by selecting a notification method according to the distance to the object. It is advisable to configure the system to use the notified information for display. In this way, safety can be improved by selecting a notification method according to the distance to the object. It is advisable to configure the system to use the notified information for AI processing. In this way, safety can be improved by having the AI select a notification method according to the distance to the object.
[0115] (1-22) The image processing method includes an image extraction step of extracting one or more image regions from a captured celestial spherical image, and an image conversion step of converting the extracted image regions into a planar image, wherein in the image extraction step, the image regions are extracted such that at least a portion of the celestial spherical image overlaps in the planar image.
[0116] In this way, it is possible to determine whether or not to notify the detection result of an object based on at least one of the following: the status of the work vehicle, the presence or absence of the object, and the location of the object. It is preferable to configure the system to use the extracted data for image processing. In this way, unnecessary warnings can be suppressed by controlling the notification according to the status of the work vehicle and the location of the object.
[0117] It is advisable to configure the system to use the extracted data for display. This would allow for notification control that takes into account the status of the work vehicle and the location of the target object, thereby reducing unnecessary displays. Alternatively, it is advisable to configure the system to use the extracted data for AI processing. This would allow the AI to analyze the status of the work vehicle and the location of the target object, and automatically adjust the optimal notification strategy.
[0118] (2-1) The system may be provided with a function to obtain an image of a predetermined region (referred to as a converted region) (referred to as a converted region image) obtained by converting an image of a predetermined region (referred to as a pre-conversion overlapping region or extracted region) in a captured image in which a part of the region overlaps (referred to as a pre-conversion overlapping region or extracted region) so that the overlapping region does not overlap, and to perform a predetermined process using at least a part of the region (referred to as a post-conversion region corresponding to the overlapping portion) in the post-conversion region image.
[0119] The transformation should be performed using at least one of the following methods: projection, mapping, function, transformation table, remapping, etc. For image transformations from one region to another, known methods should be used. For example, an IC chip with a dewarp function can be used. For example, a configuration where the dewarp function is implemented on a computer using a program (library, etc.) can also be used. A combination of these methods is also acceptable. For example, the transformation should use methods such as [examples omitted].
[0120] It is best to use a camera for shooting. A camera employing a CMOS sensor as its image sensor is recommended. The captured image should have a aspect ratio of 16:9 or 4:3. The captured image should be similar to that of a surveillance camera or dashcam. The captured image should be wide-angle. In particular, the image should not encompass the entire imaging area of the image sensor within the image circle (partially outside the image circle, such as black areas). The captured image should be a circular fisheye image. A hemispherical image is particularly recommended.
[0121] The extracted image should be neither a square nor a rectangle. The extracted image should not be a quadrilateral. The extracted image should have curved edges. There may be only one extracted region, and A. the overlapping region may be the area where parts of one extracted region overlap (for example, if the extracted region is the range from 0 to 400 degrees of a circle, the overlapping region will be the 40 degrees beyond 360 degrees). Multiple such regions may be extracted. Alternatively, B. there may be two or more extracted regions, and parts of each region may overlap.
[0122] Multiple such regions may be extracted. Both regions A and B may be extracted. Multiple regions of each region A and B may be extracted. For example, a system may be provided that has the function of converting one or more image regions in a captured image that partially overlap (regions of n or less different shapes when the number of regions is p (referred to as the extracted image region group)) into images of one or more image regions (regions of p or less different shapes when the number of regions is p (referred to as the converted image region group)) which are projected onto non-overlapping regions of at least one of the n different shapes or sizes, and in which the overlapping image regions are different, and then performing processing using the converted images.
[0123] For example, if the extraction region (B) is set to two or more (multiple) regions, and the extraction is performed so that parts of each region overlap, then a first image and a second image are extracted from the camera image as detection images for detecting an object. The region from which the first image is extracted and the region from which the second image is extracted are set to overlap in at least a portion.
[0124] The regions from which the first image is extracted and the regions from which the second image is extracted are arc-shaped regions. When the two regions are combined, they form a ring-shaped region. The regions from which the first image is extracted and the regions from which the second image is extracted should be configured so that their circumferential boundaries overlap. The first and second images should be configured to be displayable. The first and second images should be configured to be displayable in real time. The first and second images should be configured to be displayed after being converted into a rectangular shape. As a predetermined process, if an object is detected, it is good practice to include a process that allows an image suggesting the region corresponding to the object to be displayed superimposed on the first and second images.
[0125] The extracted region and the transformed region should be regions of different shapes. For example, if the system is applied to a hemispherical camera with AI object detection functionality, the camera outputs an image as a round fisheye image because it is hemispherical. This image can then be "transformed" to flatten it, and then human detection can be performed as a predetermined process. For example, the donut-shaped area around the fisheye can be flattened (e.g., made rectangular).
[0126] Since 16:9 and 4:3 aspect ratios are typically easier to process, the fisheye lens cuts off the donut-shaped area around the edge, dividing the image into upper and lower regions, and then flattens each region (for example, into a rectangle) through a "transformation." However, if the image is divided precisely at the top and bottom, objects at the dividing line will be split in half and become smaller, making object detection impossible. Therefore, it is best to create a duplicated area when creating the flattened image. This will allow for more reliable object detection. Also, for example, when flattening a round image, it is best to process it so that it is displayed as a rectangle.
[0127] The predetermined processing may be performed using a portion of the region corresponding to the overlapping portion in the converted region image (converted region corresponding to the overlapping region), but it is particularly preferable to perform the processing using the entire region corresponding to the overlapping portion in the converted region image (converted region corresponding to the overlapping region).
[0128] The predetermined processing may include image recognition processing. The predetermined processing may include outputting video that includes at least a portion of the converted region corresponding to overlapping regions. The output processing may include outputting for display on a display means. The output processing may include outputting via signal lines and communication (wired or wireless communication). The processing may include transmitting the output video.
[0129] The output process should ideally perform real-time video output. Furthermore, the output process should include a recording process for saving the video to a recording medium such as an SD card. The predetermined process may be any one of these, but it is preferable to combine these processes. For example, the recognition process could include a predetermined overlay display (e.g., a drawing enclosing an object) to indicate object detection, and the system could also output the video of this display.
[0130] Furthermore, as a predetermined process, it is advisable to perform a process in the recognition process to indicate that an object has been detected by displaying a predetermined overlay (for example, a drawing that encloses the object in a rectangle), and to record the resulting image on a recording medium. The display can be on a PC screen or on the display means of other devices (for example, in-vehicle equipment) (for example, an LCD).
[0131] The overlapping region before conversion may be identified, for example, according to the content of the predetermined process. For example, numerical settings such as the pixels to be overlapped may be used. In particular, it is good to identify the region in a way that the result of the predetermined process is better than before. For example, if the predetermined process is configured to detect the presence or absence and movement of an object using image recognition, it is good to configure the region to be determined according to the type of object. For example, if the predetermined process is configured to detect the presence or absence and movement of an object using image recognition, it is good to configure the region to be determined according to the size of the object in the image.
[0132] The pre-conversion overlapping region should be determined according to the content of the image near the boundary (for example, the size of objects that may be captured, the types of objects that may be captured, etc.). Alternatively, the pre-conversion overlapping region should be determined as a predetermined process according to the range detected by the object detection process (for example, the range captured by the camera).
[0133] The pre-conversion overlap area should be configured to be modifiable. The pre-conversion overlap area should be configured to be user-configurable. For example, its range should be set according to user actions. For example, its range should be moved according to user actions. For example, its range should be enlarged or reduced according to user actions.
[0134] "To be a configuration that is determined" could also mean a configuration that is determined in response to user actions. "To be a configuration that is determined" could also mean a configuration that is determined automatically. What was written as "to be a configuration that is determined" could also be written as "to be determined in advance."
[0135] (2-2) The predetermined processing may be performed using both the region corresponding to the overlapping portion in the converted region image (converted region corresponding to the overlapping region) and the region in the converted region image that does not correspond to the overlapping portion (converted region not corresponding to the overlapping region).
[0136] This approach allows for superior processing compared to conventional methods. For example, when displaying information, it becomes easier to see and understand. For example, when performing recognition processing, it becomes easier to further improve recognition performance.
[0137] (2-3) In particular, the predetermined processing is preferably performed using the region corresponding to the overlapping portion in the converted region image (converted region corresponding to the overlapping region) and the region in the converted region image that is continuous with the region corresponding to the overlapping portion in the converted region image (converted region corresponding to the overlapping region) and does not correspond to the overlapping portion.
[0138] This approach allows for superior processing due to its continuity. For example, when displaying data, it becomes easier to see and understand. For example, when performing recognition processing, it becomes easier to further improve recognition performance.
[0139] (2-4) The predetermined processing may also include a function to perform processing using an image of a predetermined region within the captured image where no part of the region overlaps (non-overlapping region before conversion). The predetermined processing may also include processing using an image of the non-overlapping region before conversion and an image of the converted region corresponding to the overlapping region.
[0140] For example, the image of the area showing the driver may be used as the image of the non-overlapping area before conversion in the captured image, and the image of the area showing the surroundings of the vehicle may be used as the image of the converted area corresponding to the overlapping area, and a predetermined process may be performed.
[0141] (2-5) The system may not have the functions of 2-1 to 2-4, or may have at least one of the functions of 2-1 to 2-4, and may have a function to record an event when it is determined that a predetermined event (e.g., some problem) has occurred and it is recognized that people have gathered in a region of a predetermined image (e.g., when the determination that people have gathered in a region of a predetermined image is made). The predetermined image may include at least one of the following: a captured image, an extracted region image, a converted region image, the "both" images from 2-2, or the "consecutive" images from 2-3.
[0142] In particular, it is good to have at least one of the following: the converted region image, the "both" image from 2-2, or the "consecutive" image from 2-3. For example, it would be good to have a function that records events when people gather to help in a scene such as when a person is involved in an accident in a factory or when a person suddenly collapses in a commercial facility. The longer the skip back (the temporal range to go back to past video footage saved at the time of location), the higher the probability that the moment of the accident will be captured, which is better.
[0143] For example, the system could be configured to record an event when it detects that an argument has started between people and people have gathered to stop it. For example, it could record an event if someone suddenly starts running. It could also record an event if it detects that someone has done something wrong (theft, molestation, etc.) and is running away.
[0144] Alternatively, an event recording could be triggered if a person with a strong sense of justice witnesses the scene and chases after the perpetrator. For example, if an incident occurs within a train station, an event recording could be made if a station employee sprints towards the scene. These detections should ideally be performed using image recognition processing on at least one of the images mentioned above. A configuration with multiple such systems would be beneficial.
[0145] When multiple systems simultaneously record events with nearby cameras, the likelihood of capturing desired footage increases. It would also be beneficial to have a function that transmits an event recording signal to other cameras, allowing those cameras to record based on that signal. When nearby cameras simultaneously record events, the likelihood of capturing desired footage increases. When the above-mentioned event recording conditions are met, such as a long period of time with many people gathered or running, it would be beneficial to have a function to cancel the event (for example, treat it as if the event never occurred). For example, image recognition could be used to understand the situation from one minute ago to the present, and if the difference in crowd size or running behavior is small, the event could be canceled.
[0146] (2-6) A system that does not have the functions of 2-1 to 2-5, or has at least one of the functions of 2-1 to 2-5, and has a camera position set to capture at least one of the captured image, extracted region image, converted region image, the "both" image of 2-2, the "consecutive" image of 2-3, the image of the non-overlapping region before conversion (these are called specific images), and has a function to identify how the target subject is captured in the specific image.
[0147] Identifying how the target subject is captured should involve a function that identifies where at least one of the people or objects is photographed. For example, a person detection function would be useful. In particular, it is good to use images of the non-overlapping region before conversion as the image to be identified. It is also good to have a function to identify the camera's position. The camera's position may be set according to user operation.
[0148] The camera position may be set automatically using image recognition or similar methods. In the case of forklifts, it would be beneficial to have a function that identifies the driver's head position from the head guard. The detection area for the driver and surrounding area will change depending on whether the camera is mounted in front of or behind the head guard.
[0149] The head guard height should be 95cm or more for counter-forklifts (sit-on forklifts) and 1.8m or more for reach forklifts (stand-up forklifts). Since they are usually made to the bare minimum, it would be good to have a camera that can be mounted in front to obtain images of the front and the driver, or mounted in the rear to obtain images of the rear and the driver, in order to identify the driver's position and issue a danger warning if a driver is present, and no warning if no driver is present.
[0150] For example, a camera could be installed in the driver's seat to detect the presence or absence of a driver. If a driver is present, the notification function could be enabled. If there is no driver, the detection function could be disabled as there is no need to detect people in the surrounding area. Derived from the above, notifications could be provided in response to combinations of engine ON / OFF status and the presence or absence of a driver.
[0151] For example, if the engine is ON but there is no driver, it is dangerous, so a warning should be issued. Also, when there is no driver, the warning for when the engine is OFF should be more prominent than the warning for when the engine is ON (because it is dangerous to have the engine ON when there is no driver).
[0152] For example, if there are people nearby, the system should have a function to detect them, and if a person enters a zone, it should have a function to provide different notification methods depending on whether the zone is a warning zone or a danger zone. For example, it should have a function to provide notification in warning zones and use light to notify in danger zones. The zone settings should be configured so that the user can set them by inputting a numerical value (distance from the vehicle).
[0153] It would be good to have a function to detect the presence or absence of a driver from the camera image. It would be good to configure the system to be able to perform person detection (e.g., driver detection) using image analysis technology. For example, it would be good to have a function to determine the presence or absence of a driver by identifying a QR code (registered trademark) attached to the driver's helmet, and to be able to detect cases where the QR code is present at the boundary of the cropped image. It would be good to have a function to change the area targeted for driver presence or absence determination depending on the camera's installation position. For example, if the camera is installed at the rear, the upper half of the image should be targeted for determination, and if the camera is installed at the front, the lower half of the image should be targeted for determination. For example, it would be good to have a function to detect the presence or absence of a driver according to the positional relationship with a specific object recognized in the image (such as a head guard, pillar, steering wheel, seat, etc.).
[0154] It would be good to have a function that detects the presence or absence of a driver from the camera images of multiple cameras. It would be good to have a function that determines the presence of a driver if a person is visible in a specific area (e.g., the upper half) of the camera image of the first camera, and also if a person is visible in a specific area (e.g., the lower half) of the camera image of the second camera. It would be good to have a function that determines the presence of a driver if the positional relationship between a specific object (e.g., a head guard) and a person is a specific relationship in the camera image of the first camera, and also if the positional relationship between a specific object (e.g., a steering wheel) and a person is a specific relationship in the camera image of the second camera.
[0155] It would be good to have a function that switches the notification function on or off depending on whether there is a driver or not. It would be good to have a function that enables the notification function depending on the combination of whether there is a driver or not and the status of the vehicle (e.g., engine ON / OFF, ignition ON / OFF). It would be good to have a function that enables the notification function depending on the combination of whether there is a driver or not and the presence or absence of people (e.g., workers) around the vehicle. It would be good to have a function that notifies in a manner depending on whether there is a driver or not. (For example, it would be good to have a function that notifies in a first manner (e.g., a buzzer) when there is a driver, and in a second manner (e.g., a warning light) when there is no driver.) For example, it would be good to have a function that notifies in a manner depending on the combination of whether there is a driver or not and the status of the vehicle.
[0156] It would be good to have a function that notifies in a manner corresponding to the combination of "presence or absence of a driver" and "presence or absence of people around the vehicle." It would be good to have a function that notifies in a manner corresponding to the "location of people around the vehicle." It would be good to have a function that leaves the system as is in the first area, which is close to the vehicle, and notifies in the second area, which is further away, if people are present in the second area. It would be good to have a function that allows the user to set the first and second areas. It would be good to have a function that notifies in a manner corresponding to two or more of the following: "presence or absence of a driver," "vehicle status," "presence or absence of people around the vehicle," and "location of people around the vehicle." It would be good to have a function that allows the user to set whether the notification function is enabled or disabled, and the notification manner.
[0157] The inventions described in (1) to (21), (1-1) to (1-22), and (2-1) to (2-6) above can be combined in any way. For example, one may combine all or part of the configuration of the invention described in (1) with at least part of the configuration of at least one of the inventions described in (2) and onward. In particular, it is preferable to create an invention that combines the invention described in (1) with at least part of the configuration of at least one of the inventions described in (2) and onward. Alternatively, one may extract any configuration from the inventions described in (1) to (2-6) and combine the extracted configurations. The applicant of this application intends to acquire rights to inventions that include these configurations. Furthermore, even if there are descriptions such as "in the case of" or "when," these are not meant to be descriptions that limit the configuration to that case or time. These are merely examples of better configurations, and the applicant intends to acquire rights to configurations that do not fall under these cases or times. Also, even if there is a sequence of descriptions, it is not limited to that order. Configurations with some parts deleted or the order rearranged are also disclosed, and the applicant intends to acquire rights to them as well. [Effects of the Invention]
[0158] According to the present invention, it is possible to provide a system that is superior to conventional systems.
[0159] Furthermore, the effects of the present invention are not limited to those described herein. Effects derived from the components disclosed in this specification and the drawings are also disclosed, and the applicant intends to obtain rights to such components through divisional applications, amendments, etc. For example, phrases such as "can do" or "is possible" in this specification are descriptions that clearly indicate the effects to be achieved, and there are components that demonstrate effects even without such descriptions. Moreover, there are effects that can be grasped by the component even without such descriptions. [Brief explanation of the drawing]
[0160] [Figure 1] Figure 1 is a diagram illustrating the configuration of an imaging system according to an embodiment of the present invention. [Figure 2] Figure 2 is a block diagram showing the electrical configuration of an imaging system according to an embodiment of the present invention. [Figure 3] Figures 3(A) to 3(E) are six-view drawings showing an example of the external configuration of an imaging device according to an embodiment of the present invention. [Figure 4] Figure 4 shows an example of an image captured by the imaging device according to an embodiment of the present invention. [Figure 5] Figures 5(A) and (B) show an example of the relationship between the installation of the imaging device and the imaging range of the imaging device according to an embodiment of the present invention. [Figure 6] Figure 6(A) shows an example of region shape transformation in a conventional system. Figure 6(A) shows an example of region shape transformation in an embodiment of the present invention. [Figure 7] Figure 7 shows a specific example of the transformation of the region shape shown in Figure 6. [Figure 8] Figure 8(A) shows a first example of an input image in an embodiment of the present invention. Figure 8(B) shows a second example of an input image. [Figure 9] Figure 9(A) shows the first example of the displayed image. Figure 9(B) shows the second example of the displayed image. [Figure 10]Figure 10(A) shows an example of the imaging range in another embodiment of the present invention. Figure 10(B) shows an example of the relationship between the image area and the output image in another embodiment of the present invention. [Figure 11] Figure 11 shows an example of the overall processing flow of the control device in an embodiment of the present invention. [Figure 12] Figure 12 shows an example of an image processing flow by a control device. [Figure 13] Figure 13 shows a second example of an image region in an embodiment of the present invention. [Figure 14] Figure 14(A) shows a third example of an image region in an embodiment of the present invention. Figure 14(B) shows an example of a display image converted from the image region shown in Figure 14(A). [Figure 15] Figure 15(A) shows a fourth example of an image region in an embodiment of the present invention. Figure 15(B) is an example of a display image in which the region shape has been transformed from Figure 15(A). Figure 15(C) is another example of a display image in which the region shape has been transformed from Figure 15(A). [Figure 16] Figure 16 is a perspective view of an example of a control device according to an embodiment of the present invention. [Figure 17] Figure 17 is a five-view drawing of the control device shown in Figure 16. [Figure 18] Figure 18 shows an example of the internal configuration of the control device shown in Figure 16. [Figure 19] Figure 19 is a block diagram showing the configuration of the human detection board connection diagram shown in Figure 18. [Figure 20] Figure 20(A) is a schematic diagram of the human detection specifications using the control device shown in Figure 16. Figure 20(B) is a schematic diagram of the human detection specifications from a front view. Figure 20(C) is a schematic diagram of the human detection specifications in the input image. [Modes for carrying out the invention]
[0161] Embodiments of the present invention will be described below with reference to the drawings. These drawings are used to illustrate the technical features that the present invention may adopt. The configuration and shape of the described apparatus are merely illustrative examples, and the present invention is not to be construed as being limited thereto. Various changes, modifications, and improvements can be made based on the knowledge of those skilled in the art, as long as they do not depart from the scope of the present invention. In the following description, the labeling using numbers such as 1st, 2nd, ... is for the purpose of identifying each element and does not define the number of elements. In addition, in the figures referenced in the following description, the scale may differ from that of the actual figures in order to make each component, each area, etc., recognizable. [1. Overall System Configuration]
[0162] Figure 1 is a diagram illustrating the configuration of the system in this embodiment. Figure 1 shows a schematic diagram of the vehicle 400 as viewed from the side. System 100 is a system mounted on the vehicle 400. In this embodiment, the vehicle 400 is a forklift, and may be, for example, an electric forklift. The vehicle 400 is not limited to a forklift, and may be, for example, a large transport vehicle with four or more wheels such as an automobile, bus, or truck, or a two-wheeled vehicle such as a motorcycle or bicycle, or other vehicle. The vehicle may be, for example, a vehicle of a means of transportation such as a train, monorail, or maglev train. In addition to vehicles, the present invention can also be used for other mobile bodies that carry people, such as ships and airplanes, and other mobile bodies, as well as non-mobile bodies.
[0163] The mast 401 is located on the front side of the vehicle body 400. The forks 402, also called claws, are the parts on which the load is placed. The forks 402 move up and down along the mast 401 in response to the operation of the lift lever by the driver (not shown) sitting in the driver's seat 403. The mast 401 also tilts forward and backward in response to the operation of the driver's tilt lever.
[0164] The steering wheel 404 is an operating device for operating the steering mechanism of the vehicle 400 and adjusting the direction of travel of the vehicle 400. The main switch 405 is an operating device for turning the overall power of the vehicle 400 on and off. The operating levers 406 are an operating device for controlling the operation of the vehicle 400, having various levers operated by the driver. The operating levers 406 include, for example, a lift lever, a tilt lever, a gear shift lever, and a forward / reverse lever. The gear shift lever is a lever for switching the gear number of the transmission. The forward / reverse lever is a lever for moving the vehicle 400 forward or backward. The pedals 407 are an operating device for controlling the operation of the vehicle 400, having various pedals operated by the driver. The pedals 407 include a clutch pedal and a brake pedal. The clutch pedal is a pedal for transmitting or disconnecting the power of the engine or other drive source to the transmission when starting, stopping, and shifting gears of the vehicle 400. The brake pedal is a pedal for operating the foot brake of the vehicle 400.
[0165] When the main switch 405 of vehicle 400 is turned on, power is supplied to the vehicle 400's electrical system, such as the electronic control unit, the drive motor, various sensors, and lighting fixtures. With power supplied to the vehicle 400's electrical system, the driver operates the control lever 406 and pedal 407, causing the vehicle 400 to perform actions such as driving or raising and lowering the forks 402. DC power is also supplied to the camera device 10 from the vehicle 400's switch-linked power terminal 410.
[0166] When the main switch 405 is turned off, no power is supplied to the electrical system of the vehicle 400, and even if the driver operates, for example, the operating lever 406 or the pedal 407, the vehicle 400 will not move or the forks 402 will not rise or fall. Also, no DC power is supplied to the system 100 from the switch-linked power terminal 410. On the other hand, DC power is constantly supplied from the vehicle 400's battery to the system 100 from the +B terminal (also called the constant power terminal) 411 of the vehicle 400, regardless of whether the main switch 26 is on or off. The control device 11 detects that the main switch 405 has been turned off and the power supply from the switch-linked power terminal 410 has stopped, and even after detecting that the power supply from the switch-linked power terminal 410 has stopped, it continues to receive power from the +B terminal 411 and remains powered on until a predetermined time has elapsed. This predetermined time may be specified by duration information stored in the storage medium 500 of the imaging device 12, which will be described later.
[0167] System 100 is equipment (also known as an aftermarket product) that is retrofitted to the vehicle 400, for example, by being purchased separately by the user. System 100 includes a control device 11, a camera 12, a GNSS receiver 13, a sensor device 14, a communication device 15, and a one-touch switch device 16. In this embodiment, the control device 11, camera 12, GNSS receiver 13, sensor device 14, communication device 15, and one-touch switch device 16 are unitized so that each has a separate housing. Therefore, the word "device" in the names of these devices may be read as "unit," etc. Also, two or more of these devices may be configured to have the same housing.
[0168] The control device 11, the imaging device 12, and the sensor device 14 are mounted on the underside of the head guard 408 of the vehicle 400. The GNSS receiver 13 and the communication device 15 are mounted on the support column 409 that supports the head guard 408. The mounting positions of each device are not limited to these positions. For example, the communication device 15 may be mounted in another position, as long as it does not hinder its communication function. The communication device 15 may be mounted, for example, on the underside or the topside of the head guard 408.
[0169] System 100 is a system that has the function of photographing the interior of the vehicle 400 and the area surrounding the vehicle 400. The imaging device 12 performs this photography. The imaging device 12 is an example of an imaging means for capturing a celestial spherical image. The imaging device 12 is a camera that captures a hemispherical area, or an area wider than a hemispherical area, and generates image data. The celestial spherical image may be, for example, a hemispherical image or a full celestial spherical image. The image area of the celestial spherical image is circular or elliptical. In this embodiment, the imaging device 12 is positioned near directly above the driver's seat 403 and is positioned so that the shooting direction is downward. The imaging device 12 has, for example, an imaging lens with a field of view (which may also be called a view area, field of view, etc.) of more than 180°. In this embodiment, the imaging device 12 is set so that the field of view θ in the vertical direction is around 210° and it can capture images of the surrounding 360° in the horizontal direction. In Figure 1, the range that the imaging device 12 can capture is shown using a dashed line. The imaging lens may be a lens such as a fisheye lens.
[0170] Figure 2 is a block diagram showing the electrical configuration of system 100. The control device 11 is physically and electrically connected to the imaging device 12, the GNSS receiver 13, the sensor device 14, the communication device 15, and the one-touch switch device 16 via cables. Each device other than the control device 11 operates by receiving power from the control device 11 via the power lines of the cables, and also outputs signals to the control device 11 via the signal lines of the cables.
[0171] In this embodiment, the control device 11 is also called the main unit and is responsible for controlling the system 100. The control device 11 includes a control unit 111, a storage unit 112, an audio output unit 113, a reader / writer 114, an image processing unit 115, and a detection unit 116. The control unit 111 controls each part of the system 100. The control unit 111 controls the operation of each part of the system 100 and supplies the power necessary for its operation to each part. The control unit 111 has a processor such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), ASIC (Application-Specific Integrated Circuit), and FPGA (Field Programmable Gate Array). The storage unit 112 is a main memory having, for example, RAM (Random Access Memory) and ROM (Read Only Memory). The control unit 111 temporarily stores the program read from the ROM of the storage unit 112 in the RAM. The RAM of the storage unit 112 provides a working area for the processor. The control unit 111 performs various functions by temporarily storing data generated during program execution in RAM and performing arithmetic processing. The control unit 111 further includes a timekeeping unit for measuring time. The timekeeping unit is, for example, a real-time clock. The timekeeping unit may be mounted on the processor's motherboard or may be externally connected to the processor.
[0172] The audio output unit 113 outputs sound. This sound may include, for example, notification sounds, background music, or voice messages. The audio output unit 113 includes, for example, an audio processing circuit and a speaker.
[0173] The reader / writer 114 is an example of a media holder that holds a storage medium 500 inserted into the system 100 through a storage medium insertion slot (not shown). The reader / writer 114 writes data to the storage medium 500 it holds and reads data from the storage medium 500. The reader / writer 114 may hold only one storage medium 500, but it may also be configured to hold two or more storage mediums 500 simultaneously. The storage medium 500 is a storage medium on which images taken by the system 100 are recorded, for example, an SD card. The SD card may be in any form, such as an SD memory card, miniSD card, or microSD card. The storage medium 500 may also store a program for a viewer (for example, a dedicated viewer) for playing back the stored images on an information display terminal such as a personal computer.
[0174] The image processing unit 115 extracts image regions from the input image received from the imaging device 12 and converts them into output images with different region shapes. Specifically, the image processing unit 115 extracts one or more image regions from the input image such that some regions overlap. The image processing unit 115 then converts the extracted one or more image regions so that they have different region shapes from the original image regions to produce an output image. The conversion to the output image includes, for example, a planarization process. The specific processing details of the image processing unit 115 will be described later.
[0175] The detection unit 116 detects pre-defined objects from the input image. The detection unit 116 may, for example, use an object detection model to detect objects from the input image. Here, the object detection model is a machine learning model trained by deep learning. The object detection model is trained using image data containing the objects to be detected, and bounding boxes and class labels annotated to that image data, as training data. The object detection model outputs a detection result when the object in the input image input to the detection unit is a pre-defined object to be detected. The specific processing details of the detection unit 116 will be described later.
[0176] The vehicle 400 is equipped with a display unit 420. The display unit 420 is a display that shows information regarding the vehicle's operating status and information regarding various controls performed by the control unit 111.
[0177] Vehicle 400 is equipped with a notification unit 430. The notification unit 430 is a module that issues various alerts. For example, when the detection unit 420 detects an object from the input image, the notification unit 430 outputs various alerts. The specific manner in which the notification unit 430 provides notifications will be described later.
[0178] The imaging device 12 captures images and generates image data obtained from the capture. The imaging device 12 includes, for example, an imaging lens, which is an example of an optical component, and an image sensor that captures light focused by the imaging lens. The image sensor is, for example, a CMOS (Complementary MOS) or a CCD (Charge Coupled Device). The image sensor captures images and outputs an image signal representing the captured image. The imaging device 12 generates a color (multicolor) image consisting of, for example, red (R), green (G), and blue (B) color components. The imaging device 12 may also include an A / D conversion circuit that converts the signal from the image sensor from analog to digital.
[0179] The GNSS receiver 13 receives signals from a GNSS (Global Navigation Satellite System). In this embodiment, the GNSS receiver 13 receives signals from GPS (Global Positioning System) satellites, which are a type of GNSS. The GNSS receiver 13 has an antenna for receiving signals from GNSS and a circuit for processing the signals received by the antenna. The control device 11 acquires position information (e.g., latitude information and longitude information) of the system 100 based on the signals received by the GNSS receiver 13. The position of the GNSS receiver 13 may be considered equivalent to at least one of the positions of the system 100, the control device 11, and the vehicle 400 or the driver in the vehicle 400.
[0180] The sensor device 14 has various sensors. For example, the sensor device 14 has an acceleration sensor and a gyroscope sensor. The acceleration sensor detects the acceleration acting on the vehicle 400. The acceleration sensor is, for example, a three-axis acceleration sensor that detects the longitudinal, lateral, and vertical acceleration of the vehicle 400. The gyroscope sensor is a sensor that detects the tilt of the vehicle 400. The acceleration sensor and gyroscope sensor may be used, for example, to estimate the position of the vehicle 400 by autonomous navigation when signals from GNSS satellites cannot be received. The acceleration sensor and gyroscope sensor may also be used, for example, to estimate the position by autonomous navigation when the GNSS receiver 13 cannot receive GNSS signals.
[0181] The communication device 15 communicates with external devices. The communication device 15 has, for example, a communication circuit and an antenna as a configuration for wireless communication. In this embodiment, the communication device 15 performs wireless communication compliant with the LTE (Long Term Evolution) standard, etc. In this case, the communication device 15 may be called, for example, an LTE unit. The communication standard is not limited to this, and the communication device 15 may perform communication compliant with the standards of mobile communication systems such as 4G and 5G. The communication device 15 is configured to allow the insertion and removal of a SIM card, and performs wireless communication based on the information recorded on the inserted SIM card. When the communication device 15 does not have a SIM card inserted, it cannot perform wireless communication, or wireless communication is restricted. The SIM card is, for example, a nano-SIM card. A SIM card is also called a subscriber identification module and is a storage medium that stores subscriber information related to a contract with a telecommunications carrier that manages and operates a mobile communication system, and authentication information used for authentication to perform communication compliant with the standards of this mobile communication system, etc. Authentication information is the information necessary to authenticate a user (for example, a subscriber), and may also be called subscriber identification information.
[0182] Furthermore, the communication device 15 may have a function to perform wireless communication via a wireless LAN (Local Area Network) such as Wi-Fi (registered trademark), or it may have a function to perform short-range wireless communication such as Bluetooth (registered trademark).
[0183] The communication device 15 is physically and electrically connected to the control device 11 via a cable. The cable is a cable conforming to a predetermined standard, having power lines and signal lines. In this embodiment, the cable is a USB cable conforming to the USB (Universal Serial Bus) standard, and more specifically, conforming to the USB Type-C standard. The communication device 15 operates by receiving power from the control device 11 via the power lines of the cable and outputs signals to the outside via the signal lines.
[0184] The one-touch switch device 16 is an operating device that is operated by the driver in emergencies such as accidents or other troubles. This operation should ideally be performed with a single touch, for example, by pressing it once. When the control device 11 determines that an operation has been performed on the one-touch switch device 16, it may use the communication device 15 to notify a predetermined recipient, or use the audio output unit 113 to output a warning sound or alarm sound. Alternatively, when the control unit 111 determines that an operation has been performed on the one-touch switch device 16, it may record an image on the storage medium 500 based on the image captured by the camera 12. This image recording may also be performed by the event recording function described later. The one-touch switch device 16 is preferably mounted in a position that is within reach of the driver in emergencies, etc., but is not easily hit by hands or feet during normal times, for example, on the right front pillar of the vehicle 400. [2. Image recording function of System 100]
[0185] System 100 has one or more of the following image recording functions. The image recording function is a function of System 100 that records images captured by the shooting device 12 as image data in a predetermined file format. The image data may be in still image format, but it is particularly preferable to use video image data. Examples of video formats include MPEG (Moving Picture Experts Group) format (e.g., MPEG2, MPEG4), but also AVI, MOV, WMV, etc. <2-1. Continuous Recording Function>
[0186] The continuous recording function (also called the continuous video recording function) is a function that continuously (i.e., constantly) records images captured by the camera 12 while the system 100 is in operation. When the control device 11 is executing the continuous recording function, it records image data captured from the start to the stop of the vehicle 400. The start of the vehicle 400 may be detected, for example, by turning on the main switch 405, and the stop of the vehicle 400 may be detected by turning off the main switch 405.
[0187] When the power is turned on after the main switch 405 is turned off, the control device 11 acquires image data from the camera 12 and writes the acquired image data to the memory in the control device 11, but it is preferable that it does not record to the storage medium 500. In this way, the power consumption of the system 100 during the period when the system 100 remains powered on after the main switch 405 is turned off is less than the power consumption during continuous recording. <2-2. Event Recording Function>
[0188] The event recording function is a function that records images captured by the system 100 in response to the occurrence of a specific event. An event is an occurrence for which images captured by the system 100 should be recorded, such as when the user performs sudden steering or braking operations while the vehicle 400 is in motion, or when the vehicle 400 approaches or collides with another object (e.g., an object or a person). The control device 11 determines that an event has occurred based on the measurements from the acceleration sensor and the gyro sensor. The conditions for determining the occurrence of an event are not limited to these. The control device 11 may also analyze the images captured by the camera 12 and determine that an event occurred when the vehicle 400 approached or collides with another object. The control device 11 may also determine that an event has occurred when a predetermined event recording button (not shown) or a one-touch switch device 16 is operated.
[0189] When the control device 11 determines that an event has occurred, it records images captured during a predetermined period before and after the event (hereinafter referred to as the "event recording period") onto the storage medium 500. The control device 11 may, for example, temporarily record images captured by the imaging device 12 in the storage unit 112 (e.g., RAM), and when it determines that an event has occurred, it may record the images from the event recording period read from the storage unit 112 onto the storage medium 500. For example, the control device 11 may create a single file containing images from 20 seconds before the event and 20 seconds after the event, for a total of 40 seconds. The event recording period is just an example and may vary depending on the type of event, and may also be changeable by the user. The control device 11 may record images consisting of multiple files onto the storage medium 500 for each event. The control device 11 may also record the time the event occurred, values measured by the acceleration sensor and gyro sensor (e.g., acceleration in each of the three axes), and position information acquired using the GNSS receiver 13, in association with the images onto the storage medium 500. [3. External configuration of the imaging device 12]
[0190] Figure 3 is a six-view drawing showing an example of the external configuration of the imaging device 12. Figure 3(A) is a front view, Figure 3(B) is a left side view, Figure 3(C) is a right side view, Figure 3(D) is a top view (plan view), and Figure 3(E) is a bottom view. The imaging device 12 comprises a housing 3 and mounting members 4 for attaching the housing 3 to a predetermined mounting position on the vehicle 400. The housing 3 constitutes the housing of the main body of the imaging device 12, which has an imaging lens 2. In Figure 3(D), the imaging range of the system 100 is shown using a dashed line. The housing 3 houses the image sensor and other components. The numbers "50", "40", and "50.3" in Figure 3(A) represent dimensions.
[0191] In this embodiment, the imaging lens 2 is a hemispherical lens. The imaging lens 2 has a configuration in which the lens is housed in a lens barrel made of die-cast aluminum. The lens barrel of the imaging lens 2 is supported by the edge 3A of the lens protrusion hole in the housing 3 and a lens holding part (not shown) located inside the housing 3. In this way, the imaging lens 2 is held in a state that protrudes somewhat from the front of the housing 3. The connector part 3C has a connector to which the connector 7A of the cable 7 is connected. The cable 7 is a cable for outputting image data showing the image captured by the imaging device 12 to the control device 11. The housing 3 has mounting parts 3D provided on its upper surface at intervals from left to right. The mounting parts 3D have screw holes that pass through from left to right.
[0192] Mounting member 4 is provided on the rear side of the housing 3 and is a bracket for attaching the housing 3 to a predetermined mounting position on the vehicle 400. The mounting position is the head guard 408 as described in Figure 1. Of the mounting member 4, the part that is attached to the mounting position is the mounting plate 4A. The mounting plate 4A is inclined at a predetermined angle with respect to the rear side of the housing 3, and this angle is configured to be changeable. An adhesive material such as double-sided tape is attached to the mounting plate 4A, and the mounting member 4 is attached to the mounting position on the vehicle 400 via the adhesive material. It is preferable that the mounting member 4 be configured so that the inclination of the mounting plate 4A can be changed while it is attached. Cable holding part 4B holds the cable 7 connected to the connector part 3C.
[0193] The mounting member 4 has a mounting portion 4C with a screw hole that penetrates from left to right inside. The mounting portion 4C is the part that is attached to the mounting portion 3D of the housing 3. The screw 5 is positioned to pass through the mounting portion 3D and the screw hole of the mounting member 4. In the first state, when the screw 5 is tightened, the angle of the mounting plate 4A with respect to the back of the housing 3 is fixed. In the second state, when the screw 5 is loosened, the angle of the mounting plate 4A with respect to the back of the housing 3 can be changed. With the mounting plate 4A attached to a predetermined mounting position on the vehicle 400, the user can adjust the shooting direction of the shooting device 12 (in other words, the direction the imaging lens 2 points) by changing the orientation of the housing 3.
[0194] The imaging device 12 can set various shooting ranges for the interior of the vehicle 400 and its surroundings by adjusting its mounting position, shooting direction, and combination with other cameras. This is described in Figures 10 to 12 of Japanese Patent Publication No. 2017-132298, so a detailed explanation is omitted in this embodiment. The imaging device 12 may be an imaging device having a display unit for displaying various information. Furthermore, the housing of the imaging device 12 does not have to be a rectangular parallelepiped; it may be cylindrical or have other shapes, for example.
[0195] Figure 4 shows an example of an image captured by the imaging device 12. For example, if the imaging device 12 generates an image that includes at least a hemispherical region, the image generated by the imaging device 12 will be like the image 121 shown in Figure 4. The image 121 has a circular image region. The distortion caused by the imaging lens (also called image distortion) is small near the center of the image 121, and the degree of distortion increases as you approach the circumferential direction in the radial direction. Near the center of the image 121, the subject is captured at a relatively large size, and as you approach the circumferential direction in the radial direction, the subject is captured at a relatively small size. In this example, the driver sitting in the driver's seat 403 and the steering wheel 404 are captured at a relatively large size, while the side and front of the vehicle 400 are captured at a relatively small size.
[0196] Figure 5 shows an example of the relationship between the installation of the camera 12 and the imaging range of the camera 12. Figure 5(A) shows an installation where a single camera, the camera 12, can capture images in the front, back, left, and right directions, making it suitable for recording the situation of the driver and forklift. This allows for not only understanding the circumstances of an accident in the event of one, but also for effectively conducting accident prevention training and education by incorporating an objective perspective. The captured image 121 shown in Figure 4 is an example of an image captured when the camera 12 is installed as shown in Figure 5(A). The installation is not limited to this, and as shown in Figure 5(B), it is also possible to install it with the area in front of the vehicle 400 as the imaging range. In this case, the imaging range can be up, down, left, and right, including the range of motion of the arm of the vehicle 400. Such installation locations are suitable, for example, when handling cargo at high places, when entering and exiting warehouses with low entrances, and on vehicles used for loading and unloading in and out of elevators.
[0197] Figure 6(A) shows an example of region shape transformation in a conventional system. Figure 6(A) shows an example of region shape transformation in an embodiment of the present invention. As shown in Figure 6(A), in a conventional system, a hemispherical image was used as the input image, two image regions were extracted, and the region shapes were transformed by a planarization process to generate output images arranged vertically. In this case, if a detection target, such as a person, was located near the boundary between the two image regions, the detection target would be cut off by the boundary during the detection process using the input image, and could not be sufficiently detected. Therefore, in the present invention, one or more image regions are extracted such that a portion of the image region overlaps, and the region shape is transformed by a planarization process. As a result, subjects photographed near the boundary can be photographed in the overlapping portion, they are not cut off by the boundary, and the visibility of the detection target can be ensured.
[0198] Figure 7 shows a specific example of the transformation of the region shape shown in Figure 6. In the example shown in Figure 7, two image regions are set on the left and right sides of the shooting range viewed from above the shooting device 12, such that they overlap and are of equal size. The size of the overlap can be arbitrarily changed, and it is also possible to have only one side. Then, the two image regions are flattened, and two output images are obtained, arranged vertically. Of these, output image A1, which is positioned above, is an image mainly of the front of the work vehicle, and output image A2, which is positioned below, is an image mainly of the rear of the work vehicle.
[0199] The size, position, and quantity of such overlapping areas can be set arbitrarily. Furthermore, when using multiple imaging devices 12, the system may be configured to allow each imaging device 12 to arbitrarily set the size, position, and quantity of overlapping areas in the input image. In a configuration of multiple imaging devices 12, some cameras may allow the size, position, and quantity of overlapping areas to be set arbitrarily, while others do not (cannot be changed from the default value).
[0200] Figure 8(A) shows a first example of an input image in an embodiment of the present invention. Figure 8(B) shows a second example of an input image. In the input image shown in Figure 8(A), the camera 12 is positioned above and slightly behind the driver's seat. In the input image shown in Figure 8(B), the camera 12 is positioned above and slightly in front of the driver's seat. The height of the head guard is usually made to the absolute minimum: Counter forklifts (sit-on forklifts): 95 cm or more / Reach forklifts (stand-up forklifts): 1.8 m or more. Therefore, the system is designed to either have a camera mounted in front to view the front and the driver, or mounted in the rear to view the rear and the driver, allowing the driver's position to be identified. If a driver is present, a danger warning is issued; if no driver is present, no warning is issued (or this can be varied by setting).
[0201] Figure 9(A) shows the first example of a displayed image. Figure 9(B) shows the second example of a displayed image. In the displayed image shown in Figure 9(A), a person is partially cut off near the upper left boundary of the screen. In the displayed image shown in Figure 9(B), by enlarging the overlapping area, the person, which is the object to be detected, is visible in the upper left of the screen, and the bounding box generated by person detection is displayed. In addition, in these displayed images, the input image is displayed on the right side along with the output images arranged vertically. Thus, the displayed image may include the input image.
[0202] Figure 10(A) shows an example of the shooting range in another embodiment of the present invention. Figure 10(B) shows an example of the relationship between the image region and the output image in another embodiment of the present invention. In the shooting range shown in Figure 10(A), an image region is set to be extracted excluding the central circular region in the right-hand plan view. In the shooting device 12 mounted on the vehicle 400, the driver's head is generally captured in the center of the captured image (input image), and when used for person detection around the vehicle 400, the central circular region has a low information density. For this reason, in system 100, this central region may be excluded, and the image region may be extracted from the annular (donut-shaped) portion shown in Figure 10B.
[0203] In the illustrated example, two image regions are extracted from the annular (donut-shaped) portion. The image processing unit 115 extracts the image regions by adjusting the ratio of the radial length of the annular region and the length of the arc located radially outside, based on the aspect ratio of the display area of the display image which is set in advance. Specifically, based on the radial length R1 of the annular region and the length R2 of the arc located radially outside, for example, if the size of the display area is 9:16, the height of one of the output images arranged vertically is set to 4.5. The image processing unit 115 may also adjust the size of the central circular region of the input image to be excluded from extraction so that the radial length R1 of the annular region corresponds to a length of 4.5 when the length of the arc located radially outside is R2 = 16. The size of the display area can be changed arbitrarily.
[0204] Figure 11 shows an example of the overall processing flow of the control device 11 in an embodiment of the present invention. First, the imaging device 12 captures an image (step S1100). Then, the control device 11 acquires the image (step S1101). After that, the image processing unit 115 of the control device 11 performs image processing (step S1000). Details of the image processing will be described later. After that, the control device 11 outputs the display image 420 obtained by the image processing to the display unit or administrator terminal (step S1102). After that, the detection unit 116 of the control device 11 detects the object (step S1103). Finally, the detection unit 116 outputs the detection result to the notification unit 430 or administrator terminal. As a result, the detection result is notified from the notification unit 430 or administrator terminal. Note that the order of each step described may be changed within a range that is not technically contradictory, and the processes may be performed simultaneously or repeatedly.
[0205] Figure 12 shows an example of an image processing flow by the control device. First, the image processing unit 115 reads the input image (step S1010). Next, the image processing unit 115 sets the overlapping parts (step S1020). Next, the image processing unit 115 extracts the image region (step S1030). Next, the image processing unit 115 performs a transformation of the region shape (step S1040). For example, a planarization process is performed as part of the transformation of the region shape. Finally, the image processing unit 115 generates the display image (step S1050).
[0206] Figure 13 shows a second example of an image region in an embodiment of the present invention. As shown in Figure 13, the image region may be a single region extracted from the entire input image. In the illustrated example, an overlapping portion with an angle θ is set. Therefore, in the output image A3 obtained from this image region, the region shape is transformed over a range that combines the entire circumference of the input image (360°) and the overlapping portion θ.
[0207] Figure 14(A) shows a third example of an image region in an embodiment of the present invention. Figure 14(B) shows an example of a display image converted from the image region shown in Figure 14(A). In the example shown in Figure 14(A), image regions of equal size are set at 120° intervals with respect to the input image. Each image region has an overlap of 15° on both sides in the circumferential direction between it and adjacent image regions. In the display image shown in Figure 14(B), output image A5, which mainly corresponds to the front, is displayed larger than output images A4 and A6, which correspond to the left and right. That is, in system 100, the size of the extracted image region and the size of the output image displayed in the display image do not need to be correlated. Similar to the size and number of image regions, the layout of the output images in the display image (position and size of each image region) can be freely set.
[0208] Figure 15(A) shows a fourth example of an image region in an embodiment of the present invention. Figure 15(B) is an example of a displayed image whose region shape has been transformed from Figure 15(A). Figure 15(C) is another example of a displayed image whose region shape has been transformed from Figure 15(A). As shown in Figure 15(A), image regions to be extracted may be set in four regions. In this case, the sizes of each image region may be different from each other.
[0209] As shown in Figure 15B, the displayed image may be laid out by arranging multiple image regions vertically and horizontally. Alternatively, as shown in Figure 15(C), a portion of the multiple image regions may be divided and displayed vertically. (Note) Further details regarding the components included in this invention are provided below. Hemispherical camera AI object detection Because it's a hemisphere, a round fisheye image is output from the CMOS sensor. This image is flattened before human detection is performed. The donut-shaped area around the fisheye lens is flattened. Typically, 16:9 and 4:3 aspect ratios are easier to process as standard (because they are standardization technologies). Dividing an image into upper and lower sections and flattening it is a common practice. If you divide the object precisely at the top and bottom section, the object at the dividing line will be split in half, become smaller, and then become undetectable. Therefore, when creating the image that is flattened vertically, we create overlapping areas. This allows for object detection. Also, make sure you can output the video at this time (to make it easier to show that an object has been detected by enclosing it in a rectangle). Furthermore, the output is a real-time video output (for example, displayed to the driver). Furthermore, the output is recorded on a recording medium such as an SD card. Also, the output video is transmitted. Additionally, when displaying images (round or flat) recorded on a recording medium, a square shape is added. The display can be either on a PC or the device's LCD screen. The following configurations may also be included.
[0210] 1. To detect an object, a first image and a second image are extracted from the camera image. The region from which the first image is extracted and the region from which the second image is extracted are set to overlap in at least part.
[0211] 2. The region from which the first image is extracted and the region from which the second image is extracted are arc-shaped. When the two regions are combined, they form a ring-shaped region. The circumferential boundary portions of the ring-shaped region from which the first image is extracted and the region from which the second image is extracted overlap. 3. Limiting the overlapping area (e.g., limiting it to numerical values such as pixels). 4. The overlapping areas can be changed. 4.1. The overlapping area is determined according to the type of object to be detected. 4-2. The overlapping area is determined according to the content of the image near the boundary (such as the size and type of objects that may be present). 4-3. The overlapping area is determined according to the detection range (the range captured by the camera). 4-4. The areas to be duplicated can be configured by the user. 5. The first and second images can be displayed (in real time). 5-1. The first and second images can be converted into rectangular shapes for display. 5-2. If an object is detected, an image indicating the area corresponding to the object can be displayed overlaid on the first and second images. The following configurations may also be included.
[0212] 1. A system equipped with a function to set the boundaries of a donut shape (or circle; the same applies below) where a shielding element (which may appear as a completely black area) is constantly present between the detected object (such as a pedestrian) and the camera image. The boundaries of the colors in the diagram below (for example, the area where a pillar is visible) are used as the boundaries. 1-1. Generate an image that completes the cuts of a donut, then unfold it into a predetermined shape (such as a rectangle) for display or image recognition.
[0213] 1-2. Make at least one of the cuts in the donut the end of the rectangle. (This is not mandatory, but) for example, divide the area into AB and DC as shown above and expand it into two rectangular areas as shown in the diagram below. 2. For example, you can combine Hattori's ideas as shown on the right side of the diagram below. The following configurations may also be included. 3. The cut in the donut should be at a position that includes a line segment that is consistently occluded across the entire vertical direction of the image. 4. It has a function that automatically sets the cuts in the donut. 5. It has a function that allows users to set the cuts in the donut based on their input.
[0214] 6. When transmitting video as frame data with an aspect ratio other than 1:1, such as Full HD, a completely black area will be created if using a circular fisheye lens. In other words, an area that is essentially meaningless (an area that is predetermined to be uniform image data) is used as an area for other data. For example, if a 16:9 video signal is transmitted, a circular fisheye image with a diameter of 9 will have a completely black area of 16-9=7 horizontally, resulting in an area of 7x9. Hattori's idea, which involves embedding (reduced) the rectangular image as described above into this completely black area for transmission, allows all of this to be sent in a single full HD stream.
[0215] 6-1. You can embed the rectangular image into a 7x9 black area by dividing it into several rows of stripes, or you can rotate it to create vertical stripes (for example, multiple vertical stripes) and embed it there.
[0216] 7. Display or use the data embedded in the completely black area described in 6 above using AI (for example, using the circular fisheye area for recording). The display resolution will be low. AI doesn't require that much resolution anyway. The following configurations may also be included. A camera is installed in the driver's seat to detect whether or not there is a driver. • If a driver is present, enable the notification function. If there is no driver, disable the detection function as there is no need to detect people in the surrounding area.
[0217] • Derived from the above, notifications will be issued in response to combinations of engine ON / OFF status and the presence or absence of a driver. For example, if the engine is ON but there is no driver, a warning will be issued because this is dangerous. Also, for example, when there is no driver, the warning for when the engine is OFF will be more prominent than the warning for when the engine is ON (based on the idea that having the engine ON without a driver is dangerous). For example, if there are people nearby, (Detection of people in the vicinity) • When a person enters a zone, the notification method will differ depending on whether the zone is a warning zone or a danger zone. For example, a warning zone will be notified, while a danger zone will be notified with light. • Zone settings are configured by the user entering a numerical value (distance from the vehicle). The following configurations may also be included. 1. The system detects the presence or absence of a driver from the camera image.
[0218] (The image analysis techniques in Ideas No. 2 and 3 above can also be used to detect people (in this case, driver detection). For example, when identifying a QR code attached to a driver's helmet to determine the presence or absence of a driver, if the QR code is located at the boundary of the cropped image.) 1-1. Depending on the camera's installation location, the area to be used for determining the presence or absence of a driver will be varied. (For example, if the camera is positioned at the back, the upper half of the image will be used for analysis; if the camera is positioned at the front, the lower half of the image will be used for analysis.) 1-2. The presence or absence of a driver is detected based on the positional relationship with specific objects (such as head guards, pillars, steering wheel, seats, etc.). 1-3. The system detects the presence or absence of a driver from camera images taken by multiple cameras.
[0219] 1-3-1. If a person is visible in a specific area (e.g., the upper half) of the camera image from the first camera, and a person is also visible in a specific area (e.g., the lower half) of the camera image from the second camera, it is determined that there is a driver present.
[0220] 1-3-2. If the positional relationship between a specific object (e.g., head guard) and a person is specific in the camera image of the first camera, and the positional relationship between a specific object (e.g., steering wheel) and a person is specific in the camera image of the second camera, then it is determined that there is a driver. 2. Enable the notification function depending on whether or not there is a driver.
[0221] 2-1. Enable the notification function according to the combination of "presence or absence of a driver" and "vehicle status" (e.g., engine ON / OFF, ignition ON / OFF). 2-2. Enable the notification function according to the combination of "presence or absence of a driver" and "presence or absence of people (e.g., workers) around the vehicle." 3. Notification will be provided in a manner appropriate to whether or not there is a driver. (For example, if there is a driver, the first method of notification (e.g., a buzzer) is used, and if there is no driver, the second method of notification (e.g., a flashing light) is used.) 4. Notification will be provided in a manner appropriate to the combination of "presence or absence of a driver" and "vehicle condition." 5. Notification will be provided in a manner appropriate to the combination of "presence or absence of a driver" and "presence or absence of people around the vehicle." (The image analysis techniques in Ideas No. 2 and 3 above can also be used to perform human detection (in this case, detection of the driver and people around the vehicle).) 6. Notification will be provided in a manner appropriate to the "location of people around the vehicle." 6-1. If a person is present in the first area, which is close to the vehicle, leave it as is in the first mode; if a person is present in the second area, which is further away, notify in the second area. 6-2. The first and second domains are user-configurable. 7. The information will be provided in a manner that corresponds to two or more of the following: "presence or absence of a driver," "vehicle status," "presence or absence of people around the vehicle," and "location of people around the vehicle." 8. The notification function can be enabled or disabled, and the notification method can be configured by the user. The following configurations may also be included.
[0222] In terms of evidentiary value in various senses, overlaying (embedding / compressing the OSD-displayed video) data into video, rather than saving it as text or binary data, has been a practice for quite some time. Commonly recorded information includes date and time, GPS coordinates, vehicle information (brake, speed, turn signal icons and numbers), and the vehicle's direction of travel. I think battery information is a given. Therefore, (1) Is there a patent somewhere that shows video recordings as an example? There are also methods to overlay various information onto the video (embedding it / compressing the video displayed on the OSD). Of course, we sometimes record it as data. The superimposed data is Date and time, GPS coordinates, vehicle information (brake, speed, turn signal) icons and numbers, vehicle direction of travel, various sensor information, vehicle malfunction information, battery information, etc. The following configurations may also be included. <When a problem occurs, the system detects when people gather and records the event.>
[0223] - When someone is involved in an accident at a factory or suddenly collapses at a commercial facility, the system should record the event by determining when people gather to help (the longer the skip back, the higher the probability that the moment of the accident will be captured, which is better). • Detects when an argument begins between people and when others gather to stop it, and records the event. <If someone suddenly starts running, record the event.> • Detects when a person has committed an illegal act (theft, molestation, etc.) and is fleeing, and records the event. - Or, the scene could trigger a strong sense of justice in someone who is driven to pursue the matter.
[0224] I've seen scenes where station staff sprint to the scene of an incident inside a train station. When cameras in the vicinity simultaneously record the event, there's a chance that some of the footage you want to capture might be preserved somewhere. [Example 0] In this embodiment, the image processing system includes a control unit 111 with the following functions.
[0225] Furthermore, in all descriptions within this specification, the preceding part of a sentence describing an effect, including phrases such as "can do" or "is possible" (e.g., "by doing" or "in order to do"), indicates that it is realized by the configuration provided by the control unit 111.
[0226] The functions of the control unit 111 can be implemented by a computer executing a program to achieve these functions. Of course, it may also be configured to be executed using hardware (a combination of known IP, etc.). (1) • Image extraction function One or more image regions are extracted from the input image (for example, a celestial image taken by the imaging device 12) such that some areas overlap. Image conversion function The extracted image region is converted into an output image with a shape different from the original region shape. • Ensuring visibility through overlapping sections By expanding the subject in the input image into the output image using overlapping regions, the subject can be clearly presented while minimizing cropping. • Applications to displays The system includes a configuration for displaying output images, and by projecting the management screen onto a display unit (such as a monitor), subjects in overlapping areas can be viewed without being cut off, thus improving work efficiency. • Celestial sphere image as input image
[0227] When a celestial camera is capturing images of the area around a work vehicle 400, it is possible to extract a specific area, such as an arc shape, from the wide-area image and convert it into, for example, a rectangular shape to improve visibility. • Transformation from arc shape to rectangular shape
[0228] By extracting an arc-shaped region from a celestial spherical image (circular shape) and unfolding it into a rectangular shape, the visibility of the subject can be improved. For example, the extracted region can be freely designed, such as using blind spots as boundaries. • Multiple region extraction and diverse size settings
[0229] It allows you to set not just one image area, but multiple image areas, and can flexibly handle cases where they overlap or differ in size and position. This further enhances the visibility of the subject in the overall output image. • Example of equirectangular transformation Converting a circular celestial sphere image to a rectangular shape yields a rectangular output image, making it easier to manage. • Real-time display and output to multiple devices
[0230] If the control unit 111 transmits the output image to the display unit in real time, the user can immediately view the wide-angle image. Furthermore, if the system is configured to output to multiple terminals, multiple personnel in different locations can simultaneously monitor and manage the system, increasing convenience.
[0231] It is preferable to have a configuration in which the control unit 111 extracts and transforms the image region including overlapping parts to generate an output image, and then outputs it to a display unit or multiple terminals. This allows for easy and efficient use of the wide-area information obtained from the spherical image surrounding the work vehicle 400. (2) In the image processing system of this embodiment, the control unit 111 has a function to generate a "display image" which is an arrangement of multiple output images. • Function to combine multiple output images
[0232] As an example, create a display image by arranging multiple output images extracted and transformed from a celestial sphere image horizontally or vertically. The configuration should allow the user to view information from multiple viewpoints simultaneously. • Output to the display unit By displaying a display image on a display unit (monitor, etc.) that arranges multiple output images generated by the control unit 111, the user can grasp a wide range of situations on a single screen. • Example of arranging vertically It is best to display the two output images side by side, one above the other. Because each video can be compared and referenced simultaneously, the overall viewability during operation and monitoring is improved.
[0233] (3) Furthermore, in this embodiment, the input image is assumed to be a spherical coordinate-based image (celestial sphere image projected onto polar coordinates) obtained from the imaging device 12 (Figure 4), and the control unit 111 has a function to extract an image region from the annular region excluding the circular region at the center of the coordinate system. ·Center removal extraction function The central part (circular region) of a celestial sphere image tends to have a low information density, so we extract only the annular portion, excluding that central area. Because information tends to concentrate in the peripheral areas of a ring-shaped region, a highly accurate image area can be obtained. Doing so will produce the following effects: • Improved information density Since the ring-shaped region excluding the central area contains a lot of useful information, the information density of the extracted image region can be increased, improving the effectiveness of analysis and display. • Reduced processing load and improved speed By excluding the unnecessary central portion from the processing target, the load of image conversion and recognition algorithms can be reduced, and the speed of real-time processing can be increased.
[0234] The control unit 111 generates a display image in which a plurality of output images are arranged, and further extracts and utilizes an annular region excluding the central portion from the spherical-based input image, thereby enabling a user to obtain multi-directional fields of view at one time and efficiently handle video with high information density.
[0235] (4) In the image processing system of the present embodiment, the control unit 111 is configured to refer to a preset aspect ratio of a display unit, and dynamically adjust the ratio of a radial direction and an arc length in a celestial sphere image (or an annular region) that is an input image to extract an image region. · Aspect ratio reference function Acquire the aspect ratio of a display unit (such as a monitor) and adjust the extraction parameters of the annular region according to the information. · Ratio adjustment function Correct the radial length and the length of the arc located on the outer side according to the aspect ratio of the display unit, and optimize and extract the image region. · Adaptation to various input image sizes Since the shape of the annular region can be automatically adjusted according to the display unit, uniform display and analysis can be performed with reduced user effort even for celestial sphere images of different sizes. · Improvement of user convenience Since there is no need to manually adjust the ratio for each image, work efficiency is improved. (5) Furthermore, the control unit 111 also has a function of setting (or changing) the range of an overlapping portion in an image region based on a specification from a user. · User specification reception function Through an input unit (such as a keyboard or a touch panel), the user can directly specify the range of the overlapping portion by numerical values or range specification. · Superimposed display function In a state where the input image is displayed on the screen, the overlapping portion specified by the user is visualized, making it easy to grasp an accurate range. · Application to image processing Since the user can freely set the overlapping portion, adjustments can be made to prevent the subject from being cut off depending on the application, or conversely, unnecessary regions can be cropped out. · Application to display By means of contrivances such as highlighting the specified overlapping portion on the display screen, the driver or administrator can intuitively grasp the range thereof. · User convenience Customizability is improved, and overlapping portion settings can be performed in accordance with various environments and work contents.
[0236] It is preferable that the control unit 111 has a function of adjusting the annular region according to the aspect ratio of the display unit and allowing the range of the overlapping portion to be freely set by user specification. This makes it possible to utilize the celestial spherical image around the vehicle 400 flexibly with high visibility. (6) In the present embodiment, the control unit 111 has a configuration capable of changing the range of the overlapping portion based on attributes of a detection target set in advance. · Detection target attribute setting function The user operates the input unit to perform a process of registering in advance the types of detection targets such as "large cargo", "people", and "other vehicles". · Adaptive overlapping portion changing function According to the registered attributes (size, shape, etc.), the overlapping range is changed so that the target object can be sufficiently captured in the image. · Handling large detection targets The range of the overlapping portion can be automatically expanded so that the target object can be clearly captured without being cut off. · Suppression of unnecessary regions A smaller overlapping range is applied to small target objects, so that image resources are not wasted.
[0237] (7) Furthermore, the control unit 111 may be provided with a function of determining in real time the attributes of a subject captured in an input image (for example, a celestial spherical image captured by the imaging device 12), and changing the range of the overlapping portion based on the result. · Subject attribute determination function Image analysis is used to estimate whether the subject corresponds to a "person", "forklift cargo", "obstacle", or the like. · Dynamic overlapping range changing function The process involves scaling the overlapping areas to ensure the subject fits neatly within the frame. • High responsiveness When the attributes of the subject change or a new subject appears, the overlapping areas can be readjusted at that moment to ensure proper reflection. • User-defined settings for each attribute Because different overlapping areas can be pre-set for each type of subject, optimization can be performed to suit the usage environment. (8) Furthermore, the image processing system of this embodiment also includes a function in which the control unit 111 changes the range of the overlapping portion based on the "position" of the subject in the input image. • Subject position detection function The system analyzes in real time whether a subject in an image is near a boundary or is moving. • Boundary-near overlap extension When the subject approaches the edge of the area, the overlapping area is dynamically expanded to adjust the image so that the subject is not cut off. Real-time tracking Even if the subject moves, the system senses its position and appropriately expands or shrinks the overlapping area, ensuring that important information is reliably captured. • Flexible settings By maintaining visibility near boundaries and adjusting overlapping areas as needed, the quality of video for monitoring and analysis is improved.
[0238] The control unit 111 determines the attributes of the detected object or subject, as well as the position of the subject, and dynamically changes the range of the overlapping area, thereby optimizing the captured video around the vehicle 400 and enabling high visibility and responsiveness depending on the purpose and situation of use. (9) In this embodiment, the image processing system includes a control unit 111 that changes the range of the overlapping portion based on the "shooting range" of the input image. • Shooting range detection and setting function The system automatically detects the shooting range from the metadata and analysis results of the input image (for example, a celestial image taken by the shooting device 12). Alternatively, the user can manually set the shooting range using input operations. • Improved versatility Even if the camera specifications (such as viewing angle) are different, the overlapping portion can be adaptively adjusted based on the shooting range, so it is easy to adapt to various camera environments. ·Adaptation to user needs If the user specifies the shooting range in advance, the overlapping portion can be optimized according to the application, and unnecessary areas can be cut off.
[0239] (10) Furthermore, in this embodiment, in addition to the image processing function of the control unit 111, the present invention includes a "detection means" for detecting a preset detection target from an input image, and has a configuration that highlights the area of the target in the display image when the detection means detects the target. ·Detection Means
[0240] An algorithm (such as AI object detection) that analyzes celestial images and planar images and recognizes designated objects such as people, vehicles 400, and materials. It may be executed by a program of the control unit 11, or may be executed by hardware. ·Object Display Function Processing is performed to superimpose emphasis frames and text labels (attribute information) on the area of the detection target. This mechanism provides the following advantages. ·Immediate recognition by the user Since the position of the object is visually specified, for example, a driver or manager of a forklift (vehicle 400) can immediately identify dangerous objects and workers. ·Utilization of attribute information When detection targets are displayed with labels such as "person", "vehicle", and "material", on-site judgment and management become smooth. ·Support for multiple output images
[0241] When the display image is composed of a plurality of output images, information is superimposed and displayed on the same object detected in each output image, so the user can grasp the situation without overlooking. It is preferable that the control unit 111 dynamically adjusts the overlapping portion with reference to the shooting range of the input image, and further performs highlighting on the display image when the detection means finds a target object. This enables the realization of an advanced system that efficiently monitors the situation around the vehicle 400 while providing necessary warnings and information.
[0242] (11) In the image processing system of this embodiment, the control unit 111 is equipped with a detection means and is configured to determine the presence of a driver of a work vehicle as a pre-set detection target object from the input image captured by the on-board camera (shooting device 12). • Worker code reading function The system uses image analysis to detect designated worker codes (e.g., QR codes) attached to helmets, work clothes, etc., worn by the drivers of vehicle 400. If "driver" is registered as a target object for detection, the detection of this code will determine that the driver is in the vehicle. • Identifying accurate driver information If worker codes and driver information are linked, it is possible to accurately determine which worker is driving which vehicle 400. • Matching with known work vehicle identification information
[0243] For example, a vehicle-specific ID and the operator code authorized to drive that vehicle are stored in advance, and the detection means of the control unit 111 compares the two to determine whether the operator is the appropriate driver. Specifically, the following flow can be envisioned. • Driver detection from input image The control unit 111 receives the video captured by the camera 12 and scans the area containing the worker code. • Reading the worker code If a QR code or similar is recognized, that code is linked to the driver's information. • Verification with work vehicle identification information If the code matches the ID of vehicle 400, it is determined that the vehicle is being operated by an authorized driver, and a notification can be sent to the system and management terminal.
[0244] By analyzing the driver code from images captured by the in-vehicle camera (camera 12), it is possible to accurately identify the worker riding in vehicle 400 and determine whether they are a driver with the appropriate authority. This significantly improves safety and management efficiency at the work site.
[0245] (12) In this embodiment, the control unit 111 is configured to dynamically set and change the "evaluation area for detecting the presence of a driver" in the input image based on the mounting position of the in-vehicle camera (shooting device 12 in Figure 4). • Camera position reference function Refer to the model and interior layout information of the 400 work vehicles to determine whether the cameras are located at the front, top, or center. • Evaluation domain setting function When the camera is positioned at the top center, the system automatically sets the evaluation area to a location where the driver is easily visible (such as the center or rear of the input image). This approach allows for accommodating differences in camera mounting positions depending on the vehicle model of the 400 series. • Improved driver detection accuracy Because it prioritizes evaluating areas that are less likely to be blind spots, it is easier to reliably detect the driver. • User convenience The system automatically selects the appropriate area, saving users the trouble of switching settings for each device.
[0246] (13) Furthermore, the detection means of the control unit 111 also has the function of changing the evaluation area for detecting the driver based on the positional relationship with a specific structure fixed to the work vehicle 400 (e.g., the mast or head guard of a forklift). • Structural positional relationship reference function Identify which directions and ranges a specific structure is obstructing, and eliminate blind spots from the shooting range of the vehicle-mounted camera (shooting device 12 in Figure 4). • Exclusion setting function for evaluation areas This process efficiently detects the driver by excluding areas that are practically invisible due to the presence of structures or areas that are significantly affected. • Reduced load by eliminating blind spots By pre-excluding areas hidden by structures, the detection method does not need to analyze unnecessary parts of the image, thus reducing the processing load. • Maintaining appropriate detection accuracy By eliminating unnecessary areas and focusing on evaluating the areas where the driver is likely to be captured, false detections and oversights can be suppressed. By appropriately adjusting the evaluation area based on camera position and structural position information, the accuracy and processing efficiency of driver detection for vehicle 400 can be improved simultaneously.
[0247] (14) In this embodiment, the control unit 111 performs detection processing on multiple input images obtained from multiple in-vehicle cameras (including the imaging device 12 in Figure 4), and determines the presence of a driver by combining the results. • Multiple camera individual analysis function The system analyzes images captured simultaneously by cameras A and B individually to detect the presence of a driver. • Results integration function The detection results from each camera are aggregated, and after determining whether the same person or location is present, the presence of the driver is ultimately determined.
[0248] In this way, even though the multiple cameras mounted on the vehicle 400 each cover different viewpoints and shooting ranges, they can be combined to simultaneously monitor a wide area, thereby improving the accuracy and reliability of driver detection. (15) Furthermore, the control unit 111 is configured to output the final detection result only when the detection results of the first camera and the second camera match within a predetermined range. ·Match judgment function Camera A and Camera B compare the positions and attributes of drivers and objects detected by each camera, and the results are adopted only if they match within a predetermined margin of error.
[0249] Specifically, the system outputs a final result when certain criteria are met, such as when the coordinates of the detected object in the input image are close together, or when the confidence level is above a certain level. This reduces false positives and increases the reliability of detection results by comparing false detections from one camera with the results from other cameras.
[0250] (16) The detection means of this embodiment also includes a function to confirm the presence of a driver when the relative position of the driver detected by the first camera and the second camera and the surrounding structure is within a predetermined range. • Relative position comparison function
[0251] The system calculates the positional relationship between the driver recognized by each camera and surrounding structures (e.g., the head guard of vehicle 400). If the relative positions between the two cameras do not differ significantly, it is determined that the same driver is in the same location.
[0252] For example, if the distance and angle to the head guard of a forklift vehicle are almost identical in both cameras, there is a very high probability that the subject is the same driver, and the system will perform a process to confirm their presence. This allows the system to utilize areas where the relative position to surrounding structures is stable, significantly reducing false detections and improving detection accuracy. By integrating the results from multiple cameras and comparing them with the relative positions of surrounding structures, it becomes possible to further improve the accuracy of driver detection within the vehicle 400.
[0253] (17) In this embodiment, the image processing system further includes a notification means, and is configured to notify the detection result of the object only when the detection means (control unit 111) detects the presence of a driver in the vehicle 400. • Driver presence detection and linkage function If it is determined that the driver is in vehicle 400, the notification system will be activated. The driver can be notified in various ways, including through an alert sound, a warning display on the screen, or a flashing warning light. • Secure communication Because the warning is issued only when the driver is clearly detected, it can reliably communicate surrounding hazards to the person actually operating the vehicle. • Alarms depending on the situation The system allows for precise operation, such as issuing a strong alert only when the vehicle 400 is being operated, and switching to a different operation when it is stopped or unoccupied. (18) The notification means may also be configured to notify the detection result in a different manner than when the presence of a driver is detected. • Notification mode switching function The system switches between modes such as using in-vehicle displays and audio alerts when a driver is present, and using exterior lights and warning lights to alert those around the vehicle when there is no driver. • Informing surrounding workers If hazardous materials are detected when the driver is not present, switching to a strong alert sent outside the vehicle makes it easier for nearby workers to notice. • Operational flexibility By changing the way warnings are issued according to the situation, such as during unmanned operation or when the engine is off, unnecessary warning sounds can be reduced while being emphasized when necessary.
[0254] (19) Furthermore, in the image processing system of this embodiment, the notification means has a function to change the manner of notification based on the positional relationship between the work vehicle 400 and the object to be detected. • Positional relationship determination function The system calculates the distance and angle between vehicle 400 and the target object, and changes the notification level according to categories such as "close," "medium distance," and "far distance." • Step-by-step warning function You can set up tiered warnings, such as stronger alerts (loud volume, red display, etc.) for closer objects, and weaker alerts or just a display for more distant objects. • Alerts based on risk level If the vehicle is very close to vehicle 400, a strong warning will be issued to draw attention, while if it is far away, a milder notification will be issued, allowing for a response that matches the actual risk level. • Effective information provision to the surrounding area The system enhances safety while reducing confusion at the scene by emphasizing warning sounds and lights at close range, and using only a milder display at longer distances.
[0255] By switching the notification system on or off depending on whether a driver is present, and further varying the intensity of the warning depending on the relative positions of the vehicle 400 and the object, it is possible to achieve both safety and convenience.
[0256] (20) In this embodiment, the notification means refers to the positional relationship between the object to be detected and the vehicle 400, and if the relationship is within a predetermined range, it notifies the detection result in the "first mode", and if it is outside the range, it notifies the detection result in the "second mode". • Range detection function The control unit 111 calculates the coordinates and distance between the object and the vehicle 400, and determines that if they are within a set threshold, they are in a proximity state, and if they are outside the threshold, they are in a distance state. • Notification mode switching function In close proximity (first mode), strong warnings such as alert sounds and alert displays are issued, while in distant proximity (second mode), relatively milder warnings or light illumination are limited to these. This allows for a stronger warning to be issued, for example, when the object is very close to vehicle 400, making it possible to draw attention appropriately and effectively.
[0257] (21) Furthermore, the image processing system is configured to change the notification mode based on at least one of the following, in addition to the notification means: the status of the work vehicle (such as driving information), the presence or absence of objects to be detected around the vehicle and their location. • Multi-condition notification control function
[0258] The control unit 111 is connected to the vehicle's CAN or other vehicle control network, and dynamically switches the warning method based on a comprehensive assessment of the vehicle's 400 driving status, engine status, presence and location of surrounding objects, etc., obtained from CAN information. • Focused monitoring of rear areas, etc. It can flexibly respond to different applications, such as highlighting particularly dangerous rear zones and issuing stronger alerts when objects are present within that range. • Linkage with operating status Safety can be efficiently enhanced by implementing a system where a strong warning is issued if vehicle 400 is in motion, and a milder warning is issued if it is stopped.
[0259] (22) The system of this embodiment also includes a method by which the computer in the control unit 111 performs an "image processing step of extracting one or more image regions such that a portion of the input image overlaps, and converting the extracted image regions into output images with different region shapes." • Image extraction process → Image conversion process
[0260] For example, each step described in (1) to (21) (extraction of the celestial sphere image, setting of overlapping areas, rectangularization, etc.) is executed as a series of steps on the computer provided in the control unit 111. By doing so, the various steps of the image processing system described above can be divided into individual program procedures or multiple modules and carried out accordingly. (23) Furthermore, it is desirable that the computer be equipped with programs to implement each of the image processing systems shown in (1) to (21). • System functions via program execution
[0261] By executing the program, the processor can perform the functions described in (1) to (21), such as image extraction and conversion including duplicate parts, detection of drivers and objects, and warning output by notification means. [Example 1]
[0262] (1-1) The image processing system of this embodiment processes celestial images (including hemispheres or full spheres) captured by, for example, the imaging device 12 shown in Figure 4. The control unit 111 shown in Figure 5 may have the following functions. • Image extraction function This function extracts one or more image regions from a celestial sphere image. Here, the extraction process is performed so that at least some of the extracted regions overlap when converted to a planar image. Image conversion function
[0263] Function to convert extracted image regions into planar images: The image extraction function of the control unit 111 extracts images in such a way that some parts of the celestial sphere image overlap, and the image conversion function of the control unit 111 then flattens them. This configuration offers the following advantages. • Applications in image processing Because the extracted and transformed planar images, including duplicates, make it easier to ensure the continuity of the subject, visibility and recognition accuracy are improved when analyzing the area around vehicle 400, for example. • Applications to displays By creating overlapping areas, when a flattened image is displayed on an in-car display or PC screen, it appears seamless and natural, making it easier for the user to understand the situation. • Applications in AI processing When planar images containing overlapping regions are subjected to AI analysis, features within the image can be captured more reliably, thereby improving the accuracy of object detection and person detection. (1-2) Furthermore, the control unit 111 has the following functions. • Multiple planar image compositing function A function that generates a display image by arranging multiple extracted and transformed planar images (a so-called image synthesis function).
[0264] Specifically, the control unit 111 extracts multiple image regions using its image extraction function, converts them into multiple planar images using its image conversion function, and then uses its multiple planar image synthesis function to create a display image by arranging them side by side. This synthesis process provides the following advantages. • Applications in image processing
[0265] By processing a display image consisting of multiple planar images side by side, a wide range of video can be efficiently analyzed. For example, by using a known AI algorithm within the control unit 111 in Figure 5, the front, rear, left, and right sides of the vehicle 400 can be recognized and monitored simultaneously. • Applications to displays Displaying multiple parallel planar images on a monitor simultaneously allows users to quickly grasp information from various angles. This is effective for safety checks of work vehicles such as forklifts. • Applications in AI processing
[0266] By having AI analyze multiple composite planar images in an integrated manner, it becomes easier to centrally manage the location and movement of objects over a wide area. Comprehensive judgment of multiple images, including overlapping areas, enables highly accurate object detection and warning notification.
[0267] As described above, using the spherical image acquired by the imaging device 12, the control unit 111 utilizes image extraction, image conversion, and multiple planar image synthesis functions to convert, synthesize, and display multiple planar images, including overlapping regions. This allows the user to obtain a more natural and wider field of view, and enables highly accurate analysis, especially in AI processing, by leveraging continuity. By implementing all methods such as extraction, transformation, and synthesis as functions encompassed by the control unit 111, image processing can be managed and controlled centrally. (1-3) The image processing system of this embodiment includes the following functions of the control unit 111 for the celestial spherical image acquired by the imaging device 12. • Image extraction function This function extracts multiple arc-shaped portions from the circular region of a celestial image, excluding the circular region at the coordinate center. This configuration makes it easier to divide the circular portion into arcs.
[0268] Specifically, the circular region corresponding to the coordinate center is removed, and the focus shifts to the outer circular region, extracting multiple arc-shaped images corresponding to specific directions within that region. The following applications are possible for these extracted arc-shaped images. • Applications in image processing By dividing the circular region into arcs, it becomes possible to perform processing limited to specific directions or ranges. This is useful when understanding the area around a vehicle (400) in each direction. • Applications to displays By selecting and arranging arc-shaped images, or by enlarging them, it's possible to highlight the direction or area that the user is interested in. • Applications in AI processing By applying AI analysis only to arc images corresponding to a specific direction, learning and object detection can be concentrated, making it easier to improve detection accuracy.
[0269] (1-4) Furthermore, the image extraction function of the control unit 111 includes a function to adjust the size and ratio of circular areas and arc-shaped portions based on the aspect ratio of the display unit of the displayed image that has been set in advance. Specifically, it is as follows: • Circular area dimension adjustment function This function adjusts the diameter of the circular area excluding the coordinate center, and the ratio of the diameter to the arc length of the arc-shaped portion, to match the aspect ratio of the display. This minimizes distortion and gaps when displaying arc-shaped images, allowing for proper layout. • Applications in image processing
[0270] By changing the size of the circular area to match the aspect ratio, the analysis processing performed by the control unit 111 can be optimized. For example, this stabilizes the processing accuracy by preventing the edges of the image from becoming extremely small. • Applications to displays By displaying arc-shaped images in a size that matches the aspect ratio of the display area, the screen layout becomes easier to view, making it easier for users to grasp the information. • Applications in AI processing Because the training data and images to be analyzed are standardized to a consistent aspect ratio, the uniformity of the images handled by the AI model increases, making it easier to improve recognition accuracy and learning efficiency.
[0271] In this way, the control unit 111 extracts arcs from the circular portion excluding the coordinate center and adjusts the dimensions of the circular area to match the aspect ratio of the display unit from the spherical image obtained by the imaging device 12 in Figure 4, thereby enabling efficient use of images in a specific direction. As a result, it becomes easier for users to focus on displaying and analyzing areas of interest during safety monitoring around the vehicle 400 and wide-area monitoring, and it also contributes to improving the accuracy of AI processing.
[0272] It would be beneficial to limit the extraction target to the circular arc portion and further adjust the circular area considering the aspect ratio of the display. This would allow for effective and efficient handling of images around the vehicle.
[0273] (1-5) In the image processing system of this embodiment, the image extraction function of the control unit 111 is configured to allow the user to specify the range of overlapping portions in the image area. • User-defined function A feature that allows users to manually adjust the extent of overlap. For example, it allows users to change the degree of overlap by manipulating numerical values or sliders on a GUI. The overlapping area defined in this way can be applied to the following uses. • Applications in image processing Because users can freely modify overlapping sections, the video footage around the forklift (vehicle 400 in Figure 1) can be optimally customized according to the situation. • Applications to displays The system can expand or contract overlapping sections according to user needs, and highlight necessary information on the screen. • Applications in AI processing By allowing users to adjust the degree of overlap, the system can optimize the image so that objects appear more connected during AI learning, which helps improve identification accuracy.
[0274] (1-6) In addition, the image extraction function of the control unit 111 may be configured to automatically adjust the range of overlapping parts based on the attributes of the detected object (e.g., person, vehicle, material, etc.) that have been set in advance. • Detected object attribute recognition function For example, if you primarily want to detect "people," you can set a larger overlap range so that people are less likely to be interrupted even near the boundary, optimizing the detection method according to the type of object being detected. • Applications in image processing When the object is large or has a complex shape, ensuring a wide overlap area according to its attributes makes reliable detection easier. • Applications to displays Because overlapping areas can be adjusted to ensure that the detected object is clearly visible, it is possible to provide easy-to-view images. • Applications in AI processing This reduces the likelihood of objects being cut off or distorted in images handled by AI, contributing to improved accuracy in training data and inference results.
[0275] (1-7) Furthermore, the image extraction function of the control unit 111 may be configured to set the range of the overlapping portion based on the attributes of the subject captured near the boundary of a pre-set image area (e.g., a driver in a vehicle, a person passing by, cargo on a forklift, etc.). • Subject recognition function near boundaries If a subject is visible near the edge or seam of an image area, the system adjusts the overlapping area according to the type of subject to ensure that the information is not interrupted. • Applications in image processing Because important subjects near the boundary (for example, the image of the driver captured by the camera 12 in Figure 4) are well covered, false detections and oversights can be reduced. • Applications to displays By appropriately setting the overlapping area so that the subject is not divided by boundaries, the information the user wants to see can be clearly presented. • Applications in AI processing During learning and inference, important subjects are less likely to be lost at boundaries, making it easier for the AI to prioritize detection and identification.
[0276] It would be beneficial to incorporate user-specified settings, attributes of the object to be detected, and overlapping range settings that take into account information near the boundaries of the subject into the image extraction function of the control unit 111. This would allow for flexible customization and optimization of images, such as those around a forklift (example of vehicle 400), enabling greater convenience and accuracy in display and AI analysis.
[0277] (1-8) In this embodiment of the image processing system, the image extraction function of the control unit 111 analyzes the position of subjects captured in the celestial spherical image and sets the range of the overlapping portion based on that position information. ·Subject position analysis function The celestial image obtained from the imaging device 12 is analyzed to determine the position coordinates of the objects (people, objects, vehicles 400, etc.) captured in the image. • Dynamic adjustment function for overlapping range Identify areas where many subjects are concentrated or areas of high importance, and set the settings to include these areas in the overlapping range. • Applications in image processing
[0278] By including the area containing important subjects within the overlapping portion of the celestial image, the target of analysis can be clarified. For example, it becomes easier to accurately capture the movements of people working around a forklift (vehicle 400 in Figure 1). • Applications to displays By displaying overlapping areas that the user wants to check (such as the direction in which the subject is concentrated), visibility is improved and oversights are reduced. • Applications in AI processing Because the overlapping area can be optimized according to the subject's position, unnecessary areas can be eliminated, allowing the AI to focus on learning and inference in important areas. (1-9) In addition, the image extraction function of the control unit 111 may include a function to set the range of the overlapping portion based on the shooting range of the celestial spherical image itself. • Shooting range detection function A process to identify the overall field of view and shooting range covered by the imaging device 12 (Figure 4). Range-based overlap setting function This process uses the above-mentioned shooting range to eliminate excessive overlap and unnecessary areas, thereby performing optimal image extraction. • Applications in image processing Because the minimum necessary overlap can be set based on the shooting range, the amount of data is suppressed and the efficiency of analysis is improved. • Applications to displays Because images can be laid out according to the shooting range when displayed on the screen, the information that users need to see can be presented concisely. • Applications in AI processing By optimizing overlap based on the shooting range, the AI's input data does not become unnecessarily wide-ranging, enabling efficient analysis by eliminating unnecessary parts.
[0279] The control unit 111 appropriately sets and adjusts overlapping areas while referring to the position of subjects in the celestial image and the shooting range itself. This allows for flexible optimization of image extraction according to the user's needs, such as when accurately capturing the surrounding environment of the vehicle 400 or when improving the efficiency of AI learning.
[0280] (1-10) The image processing system of this embodiment may further include a display unit (for example, a display unit 420 or at least one of the display units of the system 100) and be configured to display multiple display images generated by the control unit 111 simultaneously. ·Display section For example, monitors or LCD displays. This allows users to visually check the surrounding conditions of the vehicle 400 (Figure 1) and the monitoring area. It is preferable to display multiple display images (extracted, transformed, and combined by the control unit 111 from the celestial image) simultaneously on this display unit. • Applications in image processing This makes it easier to build a system that displays images from different perspectives simultaneously and then analyzes them further. • Applications to displays For example, by displaying images from a front camera and a rear camera, or an overhead view, side by side, it becomes possible to compare and judge information from multiple viewpoints. • Applications in AI processing If AI can integrate and analyze multiple displayed images, it can leverage differences in viewpoints to improve the accuracy of object recognition and location estimation.
[0281] (1-11) Furthermore, this embodiment includes a "detection means" for detecting pre-set objects from the displayed image, and when the detection means finds an object, it is provided with a function to superimpose a rectangular frame onto the displayed image on the display unit. • Detection means For example, this includes the AI object detection module mounted on the control unit 111. It takes spherical or planar images as input and recognizes pre-specified people, vehicles, and other objects. • Overlay display function The area where the detected object exists is enclosed in a rectangular frame (bounding box) and displayed on the screen. • Applications in image processing Because it allows for an intuitive understanding of the object's location, it is useful for safety management of vehicles 400 and surrounding workers. • Applications to displays By highlighting them with a rectangular frame, users can quickly recognize hazardous materials or items to be monitored. • Applications in AI processing With the detected objects clearly annotated, the AI can undergo retraining and correction, leading to further improvements in accuracy.
[0282] (1-12) In addition, in the system of this embodiment, assuming that the celestial image is captured by an on-board camera (e.g., camera 12) mounted on a work vehicle, the detection means is configured to detect the presence of a specific worker by reading the worker code. • Work vehicle camera It is mounted on vehicle 400 and captures the driver's seat and surroundings as a celestial image. • Worker code reading function For example, the system detects QR codes attached to helmets or work clothes to identify who is riding or working. • Applications in image processing The system can detect the target worker's code and accurately determine whether that individual is currently driving vehicle 400. • Applications to displays The system displays information about recognized workers on the display unit, making the situation on site visible. • Applications in AI processing Code recognition enables automatic personal identification and behavioral analysis, allowing for advanced features such as automatic recording of work time and detection of abnormal behavior.
[0283] By combining simultaneous output of multiple display images to the display unit with recognition and overlay display of object and worker codes using detection means, it is possible to improve the safety of vehicle 400, enable efficient monitoring at work sites, and perform detailed AI analysis.
[0284] (1-13) In the system of this embodiment, the detection function mounted on the control unit 111 is configured to dynamically change the "evaluation area" in the displayed image in which the presence of a worker is detected by referring to the position of the onboard camera (shooting device 12 in Figure 4). • In-vehicle camera position reference function The range and position of the evaluation area are changed according to the camera installation location (front, rear, side, etc.) within the vehicle 400 (Figure 1). • Applications in image processing By adjusting the evaluation area based on the camera's relative position, it becomes possible to accommodate blind spots and differences in image quality, thus enabling accurate detection. • Applications to displays Because information aligned with the camera position can be highlighted on the user screen, optimized monitoring can be achieved, such as from the driver's seat view of the 400 vehicle. • Applications in AI processing By appropriately setting evaluation regions that take camera position into account during the learning and inference stages, the accuracy of video analysis tends to improve.
[0285] (1-14) The control unit 111 also has a detection function that changes the evaluation area for detecting the presence of a worker in the displayed image based on the positional relationship with a specific structure (e.g., the mast or head guard of a forklift). • Structural positional relationship reference function The evaluation area is adjusted based on the distance and angle to the mast, head guard, or cargo handling equipment fixed to the vehicle 400. • Applications in image processing By distinguishing between areas easily obscured by structures and areas with little obstruction, and changing the evaluation area accordingly, optimal detection can be performed according to the environment. • Applications to displays By highlighting areas that are difficult to see due to overlapping structures, or conversely, by removing unnecessary backgrounds, it is possible to provide a display that is easy for drivers to check. • Applications in AI processing Because the AI, having learned the positional relationship with specific structures, performs analysis while understanding the context of the work site, it is expected to reduce misjudgments and improve recognition accuracy.
[0286] (1-15) Furthermore, in this embodiment, if there are multiple in-vehicle cameras, detection processing is performed individually on the display images acquired by each camera, and the results are integrated. • Multiple camera individual detection function The control unit 111 separately detects objects and people in the images captured by cameras A and B, and then combines the results to make a judgment. • Applications in image processing Because it detects the same worker from multiple perspectives, it is easier to obtain highly accurate consensus results. It also has the effect of filling in blind spots. • Applications to displays By displaying integrated results, users can grasp accurate information that combines multiple perspectives. • Applications in AI processing By fusing image data from different angles, it is possible to achieve comprehensive recognition without redundancy or omissions.
[0287] (1-16) Finally, the detection function of the control unit 111 can be modified to adopt an algorithm that, for example, only adopts a detection result if the detection results of the first camera and the second camera match within a certain range. ·Result match judgment function If the coordinates, confidence level, and area size indicated by the two cameras match within a predetermined range, the system determines that the reliability is high and adopts it as the final "detection detected" result. • Applications in image processing Because they mutually correct false detections from different cameras, the false detection rate can be reduced. • Applications to displays Users can view less error-prone information by only displaying detection results where matches have been confirmed across multiple cameras. • Applications in AI processing Using only highly accurate data for training and inference reduces false positives in AI and makes it easier to achieve stable performance.
[0288] By incorporating dynamic changes to the evaluation area that take into account the position of the in-vehicle camera (shooting device 12) and its relationship to structures, and methods for integrating detection results across multiple cameras, the accuracy of detecting workers around the vehicle 400 can be further improved, strengthening the reliability of user display and AI analysis.
[0289] (1-17) In the system of this embodiment, the detection function configured in the control unit 111 compares the relative positions of the worker and surrounding structures based on the display images of the first camera (shooting device 12 in Figure 4) and the second camera, and employs a method of detecting the presence of the worker when the degree of agreement is within a predetermined range. • Relative position comparison function
[0290] The system acquires the position of the worker detected by each camera and the positional relationship between the worker and surrounding structures (such as masts and head guards). If the relative positions of two cameras are similar, it is determined that the same worker is present. • Applications in image processing Even if detection is performed from different viewpoints using each camera, highly reliable results can be obtained if the relative positions are consistent. • Applications to displays By displaying the detection results of agreed-upon workers on the user screen, false alarms can be minimized, and highly accurate information can be provided. • Applications in AI processing By training the system to learn the relative positions between viewpoints, the coordinated analysis of multiple cameras can be enhanced, improving the accuracy of worker detection.
[0291] (1-18) Furthermore, in this embodiment, the detection results regarding objects (e.g., other vehicles, obstacles, people, etc.) detected by the image recognition (an example of a detection means) of the control unit 111 are communicated to the user and the surroundings by a notification unit (at least one of the voice output unit 113 or the notification unit 430). In particular, it is preferable to control the control unit 111 to notify the detection results of objects when it detects that a worker is present on the work vehicle 400. • Hochi Department It features a variety of notification methods, including buzzers, alarm sounds, display indicators, and indicator lights. • Applications in image processing By issuing alerts only when a worker is on board, false alarms and unnecessary warnings are reduced, and warnings can be given at the appropriate time. • Applications to displays Assuming an operator is present, warning messages and target frames can be displayed on the screen to immediately alert the operator. • Applications in AI processing Because the system only notifies the operator of the target object when they are definitely operating the vehicle or are in the vicinity, it facilitates the implementation of situation-appropriate automatic control and alert output.
[0292] (1-19) If the presence of a worker is not detected in the work vehicle 400, the system may be configured to notify the detection result of the object in a manner different from that described above. For example, if there is no passenger, a warning may be given in a different way, such as sounding an alarm of a different volume or turning on an external lamp. • Notification mode switching function The warning sound, warning light, and display content emitted by the notification unit are switched depending on whether there are workers present or not. • Applications in image processing It's possible to differentiate between warnings, such as a minor warning when no workers are present and a major warning when workers are present. • Applications to displays When there are no passengers, the system can be used to highlight warnings directed outside the vehicle, allowing for flexible operation such as informing drivers of the risks of unmanned driving. • Applications in AI processing By incorporating an algorithm that issues a different warning when no worker is present, the AI can more easily select the optimal action for the given situation.
[0293] Combining worker detection based on relative position matching between multiple cameras, and switching notification modes depending on the presence or absence of workers, would be beneficial. This would further enhance safety management around vehicle 400 and improve work efficiency.
[0294] (1-20) In this embodiment, when the control unit 111 transmits information about an object detected by the detection means to the surroundings via the notification unit, it is configured to change the notification mode based on the positional relationship between the work vehicle 400 and the object. • Positional relationship reference function The system compares the coordinates of the object with the position information of vehicle 400 to calculate relationships such as distance and direction. • Notification mode switching function Depending on the relative position, the notification sound, display format, and warning light illumination pattern are switched. • Applications in image processing By adjusting the content and intensity of notifications according to the relative location, it is possible to provide clearer notifications, especially in cases of high risk. • Applications to displays It makes it easier to present appropriate information to the user, such as visually indicating the distance to an object on the display. • Applications in AI processing A system can be built in which AI continuously learns the relative position to an object and automatically adjusts the optimal notification pattern.
[0295] (1-21) Furthermore, in this embodiment, it is preferable to introduce a stepwise mechanism in which the control unit 111 provides notification in a first manner (for example, a strong warning or a loud sound) if the above positional relationship is within a predetermined range, and provides notification in a second manner (a milder warning or a gentler sound) if it is outside the range. • Distance threshold-based notification function For example, you could set thresholds such as "Phase 1 if within 2 meters" and "Phase 2 if exceeding 2 meters." • Applications in image processing The closer the object, the higher the priority given to notification, ensuring safety while avoiding unnecessary warnings. • Applications to displays Intuitive visualization is possible, such as displaying red on the screen when an object enters a dangerous area, and yellow otherwise. • Applications in AI processing The AI can continuously learn a logic that estimates the distance and orientation of an object and selects a notification method based on whether it exceeds a predetermined threshold.
[0296] (1-22) In this embodiment, the control unit 111 employs an image processing method that includes an image extraction step of extracting one or more image regions from a celestial spherical image acquired by the imaging device 12, and an image conversion step of converting those image regions into a planar image, as the processing steps (or combination of hardware IPs) of the control unit 111. Here, the image extraction step is characterized in that it extracts regions such that they overlap with parts of the celestial spherical image in the planar image. ·Image extraction process The celestial image is analyzed, and multiple regions, including overlapping ones, are extracted. • Image conversion process The extracted region is unfolded into a planar image, and processing is performed to improve visibility and analysis efficiency.
[0297] In conjunction with this, it is preferable to configure the system to control whether or not to notify the detection result of the object by referring to the status of the work vehicle 400 (whether or not it is in operation), the position of the object (degree of proximity), etc. • Applications in image processing By suppressing or strengthening notifications according to the operating status of the 400 work vehicles and their relative position to the target object, unnecessary alarms are reduced, and necessary warnings can be issued accurately. • Applications to displays The system can optimize the information users see by omitting or changing warning displays on the screen in conjunction with the operational status. • Applications in AI processing By analyzing vehicle driving information and the location of objects using AI, the system can automatically adjust the notification strategy and provide optimized safety measures.
[0298] By combining methods such as switching notification modes based on positional thresholds and extracting and transforming celestial images with overlapping regions, hazard prediction and situational awareness around vehicle 400 become more accurate and flexible. If integration with AI is also considered, a system can be built that minimizes the burden on operators while achieving high safety. (1-23) Furthermore, for (1-1) through (1-23), it is preferable to have a media recording function that records the output or internally generated information onto the storage medium 500.
[0299] Furthermore, the contents described in the "Application to ○○" section above should be implemented as the execution (processing) of a program by the computer provided in the control unit 111. Other configurations described should also be implemented as the execution (processing) of a program by the computer provided in the control unit 111. [Example 2]
[0300] (2-1) The system of this embodiment may include a function to obtain an image of a predetermined region (hereinafter referred to as the "converted region") (hereinafter referred to as the "converted region image") obtained by converting an image of a predetermined region (hereinafter referred to as the "overlapping region before conversion" or "extracted region") in an image captured by the imaging device 12 (camera) shown in Figure 4, where a part of the region overlaps (hereinafter referred to as the "overlapping region before conversion"), and then perform a predetermined process using at least a part of the region corresponding to the overlapping part in the converted region image (hereinafter referred to as the "overlapping region corresponding converted region").
[0301] In this system, transformations may be performed using at least one of the following methods: projection, mapping, function, transformation table, remapping, etc. Image transformations from one region to another can be performed using known methods. For example, this may be done on a computer using an IC chip with fisheye distortion correction (dewarp) functionality or a program (library, etc.) that implements the dewarp functionality. A combination of these may also be used. The transformation may also be configured to use, for example, equiangle projection.
[0302] The image capture should be performed using the imaging device 12 (camera) shown in Figure 4. A camera employing a CMOS sensor as its image sensor is recommended. The captured image may be in a 16:9 or 4:3 aspect ratio, and a wide-angle image, such as those obtained from a surveillance camera or dashcam (which can also be configured as imaging device 12 or as part of this system), is particularly desirable. In particular, a configuration that captures an area where the entire imaging area of the image sensor does not fit within the image circle (e.g., a black area outside the image circle) is also acceptable. The captured image may be a circular fisheye image (especially a hemispherical image). Furthermore, the extracted image may be a shape that is neither square nor rectangular, or an image shape other than a rectangle with curved edges.
[0303] The extraction region may be set to a single region. For example, A. The overlapping region may be a region where a part of one extraction region overlaps (for example, if the extraction region is set to the range from 0 to 400 degrees of a circle, this corresponds to the region where 40 degrees beyond 360 degrees overlaps). Multiple such regions may be extracted. Also, B. Two or more extraction regions may be set, and parts of each extraction region may overlap. Both regions A and B may be extracted, and a configuration may be used in which multiple A and B are extracted. For example, it is preferable to have a function that converts one or more image regions in a captured image that partially overlap (regions of n or less types of shapes when the number of regions is n, hereinafter referred to as the "extracted image region group") into one or more image regions where the overlapping region is projected onto a different region by at least one of n different shapes or sizes (regions of p or less types of shapes when the number of regions is p, hereinafter referred to as the "converted image region group"), and performs processing using the converted image.
[0304] For example, as an example of extracting multiple extraction regions as described in B. above, where at least a portion of each region overlaps, consider the case where a first image and a second image are extracted from a camera image as detection images for detecting an object. The regions from which the first image is extracted and the regions from which the second image is extracted are set to overlap at least a portion of each other. These first and second images are extracted as arc-shaped regions, and when combined, they form a ring-shaped region. Furthermore, it is sufficient to design the circumferential boundary portions of the ring-shaped region to overlap. The first and second images can be configured to be displayed (including real-time display), and both images may be converted into rectangular shapes for display. In addition, when an object is detected by a predetermined process (such as object detection processing), it is convenient to configure the system to display an indication of the object by overlaying it on the first or second image.
[0305] The extracted region and the converted region should be set to have different shapes. For example, the shooting device 12 could be a "hemispherical camera with AI object detection function," which, being hemispherical, captures a round fisheye image. After shooting, the image could be flattened as a "conversion" process, and predetermined processing such as human detection could be performed. For example, the area around the fisheye (donut-shaped part) could be flattened (e.g., made rectangular). Generally, aspect ratios such as 16:9 and 4:3 are easy to process using standard methods, so the donut-shaped area around the fisheye could be divided, and the top and bottom could be flattened and processed using "conversion." However, if the image is precisely divided vertically, objects captured at the boundary may be split, reducing recognition accuracy. Therefore, intentionally creating an overlapping region (doubled part) when creating the vertically flattened image can more reliably detect objects. Also, when flattening and displaying a round image, it may be displayed as a rectangular frame.
[0306] Thus, the predetermined processing may use only a portion of the overlapping area (converted area corresponding to the overlapping area) in the converted area image, but it may also be configured to use the entire area. The predetermined processing is preferably an image recognition process (such as object detection or person detection). Alternatively, it may be a process that outputs video including at least a portion of the converted area corresponding to the overlapping area. The output processing may include a configuration that utilizes the control unit 111 and the storage unit 112, etc., to output real-time video to the display unit 420, transmit via communication (wired / wireless), or record to a storage medium 500 such as an SD card. By combining these processes, it is highly convenient to be able to, for example, perform an overlay display (such as drawing a rectangle around an object) when an object is detected in the recognition process, output the video, or store it simultaneously.
[0307] Furthermore, the pre-conversion overlapping region (extracted region) may be configured to be specified according to the content of the predetermined processing. For example, the number of pixels to overlap may be set numerically. In particular, to improve the results of image recognition processing, the overlapping region may be determined according to the type of object (size and shape). To avoid objects near the boundary being cut off, extending the pre-conversion overlapping region within a certain range may improve detection accuracy.
[0308] The pre-conversion overlap area may be determined based on a predetermined process, such as the imaging range of the imaging device 12 (e.g., object detection range). Furthermore, the pre-conversion overlap area may be made changeable according to user operation, for example, allowing the user to move, enlarge, or reduce this overlap area from a settings screen. This "determination configuration" may be performed by user operation, automatically determined, or predetermined as a fixed value.
[0309] (2-2) In the system of this embodiment, a predetermined process may be performed that uses both the overlapping portion of the "converted region image" obtained by converting the image captured by the imaging device 12 in Figure 4 (page 4) (hereinafter referred to as the "converted region corresponding to the overlapping region") and the region other than the overlapping portion of the same converted region image (hereinafter referred to as the "converted region not corresponding to the overlapping region"). Doing so makes it possible to perform processing that is superior to the conventional method. For example, when displaying the output image, including both regions makes the image easier to see and understand. In recognition processing, it is also easier to further improve the accuracy of distinguishing objects and people by utilizing both regions.
[0310] (2-3) In particular, in the predetermined processing described above, it is preferable to use a configuration that combines the region corresponding to the overlapping portion in the converted region image (converted region corresponding to the overlapping region) with a non-overlapping region converted region that is adjacent to it. Providing continuity enables superior processing, improves readability in the case of display, and can be expected to further improve accuracy in the case of recognition processing.
[0311] For example, when the control unit 111 performs image synthesis or rendering processing, the converted area with overlapping region support and its contiguous portion (converted area without overlapping region support) can be displayed and output together, or recorded in the storage unit 112 or storage medium 500, making it easier to distinguish the driver and surrounding objects. This is also effective for ensuring safety during work around forklifts (for example, vehicle 400 in Figure 1 (page 1)).
[0312] (2-4) Furthermore, the predetermined processing may include a function that also uses images of a predetermined region within the captured image that does not partially overlap (hereinafter referred to as the "non-overlapping region before conversion"). Specifically, a configuration in which the image of the non-overlapping region before conversion and the image of the converted region corresponding to the overlapping region are combined and processed is conceivable.
[0313] For example, one method would be to use the "area showing the driver" as the non-overlapping area before conversion in the captured image, and the "area showing the area around the vehicle 400 (e.g., a forklift)" as the image of the converted area corresponding to the overlapping area. This allows for the combined processing of video including the driver (captured by the camera 12) and video showing the surrounding situation. For example, it is also possible to configure the system to simultaneously analyze and recognize the driver's state and approaching objects in the surroundings, and then transmit the results via the communication device 15, or output notifications and warnings via the control unit 111.
[0314] In this way, by acquiring wide-angle and fisheye images with the camera 12, and then converting them into overlapping and non-overlapping regions, before combining, displaying, recognizing, and recording them, it becomes easier to utilize both driver footage and vehicle surrounding footage. This provides a variety of benefits, such as improved safety in the work environment and more efficient support for workers.
[0315] (2-5) The system of this embodiment may be configured without the functions of (2-1) to (2-4), or with at least one of these functions, and may be configured to record an event when it is determined that a predetermined event (for example, some trouble or abnormal situation) has occurred and it is recognized that people have gathered in an area within a predetermined image (for example, when it is determined that people have gathered in an area within a predetermined image).
[0316] Here, the "predetermined image" should include at least one of the following: the raw image captured by the camera 12, the extracted region image extracted based on it, the converted region image, the "both" images mentioned in (2-2), or the "consecutive" images mentioned in (2-3). In particular, using images that have undergone processing of overlapping regions, such as the converted region image, the "both" images in (2-2), or the "consecutive" images in (2-3), makes it easier to grasp the surrounding conditions of the vehicle and the movements of people in more detail.
[0317] For example, in a factory setting, the system could start recording events when it detects when people gather to help someone involved in an accident, or when someone suddenly collapses in a commercial facility. Setting a longer skip-back time (saving past footage) increases the likelihood of capturing the moment of the accident, making it more useful. Recording can also be initiated when an argument breaks out and people gather, or when someone suddenly starts running. Furthermore, the system could also trigger event recording when it detects a person fleeing after a suspicious act such as theft or molestation, and when people are seen chasing them. Another example would be a system that records an event when it detects a station employee sprinting at full speed in the event of an incident in a train station. These detections can be performed using image recognition processing targeting the extracted region image or the converted region image (either overlapping or non-overlapping regions) as described above.
[0318] In addition, by configuring the system to include multiple units, and having multiple nearby imaging devices 12 (Figure 4 (page 4)) simultaneously record events, the likelihood of one of the cameras capturing the decisive moment increases. When the conditions for event recording are met in one camera, a function may be provided to transmit that signal to other cameras, allowing them to start recording as well.
[0319] Furthermore, if a situation persists where many people gather or run for an extended period, it may be possible to implement a function to temporarily cancel the above event recording conditions. For example, by using image recognition to understand the situation from one minute prior, it would be possible to implement control so that an event is not considered to have occurred in situations where congestion is constant, such as when the difference in the number of people or running behavior is small.
[0320] In this way, by introducing functions for detecting crowds and behavior using designated images, and for recording events, a safety management system that can respond quickly to unforeseen accidents and incidents can be realized. Another feature is that it enables flexible operation according to the situation on site by adding a configuration that links multiple cameras and a function that excludes prolonged crowded conditions.
[0321] (2-6) In the system of this embodiment, in configurations that do not have the functions of (2-1) to (2-5), or configurations that have at least one of the functions, it is preferable to set the position of the camera (e.g., shooting device 12) in advance so that it is possible to identify how the target subject is captured from the captured image, extracted region image, converted region image, the "both" images in (2-2), the "continuous" images in (2-3), and the pre-conversion non-overlapping region image (collectively referred to as the "specific image"). As a method for identifying the presence of the target subject, one can consider having a so-called person detection function that determines where at least one of a person or an object is captured. In particular, if the pre-conversion non-overlapping region image is used as the specific image, it is easier to separately manage the area in which the driver or object is reliably captured.
[0322] By incorporating a function to determine the camera's position, it becomes possible to determine the camera's position according to user operation or to automatically adjust the camera's position using image recognition. For example, if vehicle 400 is a forklift, it is effective to have a function to set the camera position based on the driver's head position from around the head guard (head guard 408). The detection area for the driver and surroundings changes depending on whether the camera is mounted in front of or behind the head guard. The height of the head guard is generally set to 95 cm or more for counterbalanced forklifts (forklifts that are operated while seated) and 1.8 m or more for reach forklifts (forklifts that are operated while standing), so the configuration can be configured to select whether to obtain forward-facing images + images of the driver or rear-facing images + images of the driver, depending on the camera's mounting location. This can also be applied to control systems that identify the driver's position and issue a danger or warning if the driver is present, and no warning if the driver is not present.
[0323] For example, a system could be constructed to install a camera 12 in the driver's seat to detect the presence or absence of a driver. This would allow for selective control, such as activating the notification function only when a driver is present, and disabling the function if there is no need to detect people in the surrounding area when no driver is present. Furthermore, the notification behavior could be changed depending on the combination of engine ON / OFF status and the presence or absence of a driver. For example, a warning could be issued if the engine is ON but the driver is absent, as this indicates a high level of danger, while the level of caution could be lowered if the engine is OFF.
[0324] Furthermore, if there are people nearby, the system may be equipped with a function to detect their location and issue different alerts depending on whether they enter a warning zone or a danger zone. For example, warning zones could be alerted by voice or a buzzer, while danger zones could be alerted more strongly by lights such as flashing lights. It would be convenient to allow users to flexibly determine the zone settings by numerically inputting the distance from the vehicle (400).
[0325] Furthermore, by incorporating a function to automatically detect the presence or absence of a driver from the camera image, it is possible to implement person detection (such as driver detection) using image analysis technology. For example, a configuration that can detect a QR code attached to the driver's helmet and recognize it without false detection even if the QR code is visible at the edge of the cropped image is also a good idea. A configuration that switches the target area for determining the presence or absence of a driver (e.g., upper half / lower half of the image) depending on the camera's installation position is also effective. It is also possible to incorporate an algorithm that determines whether or not a driver is present by looking at the relative position of a person and a specific object (such as a head guard, pillar, steering wheel, or seat).
[0326] If multiple cameras 12 are installed, a mechanism can be considered to determine the presence or absence of a driver from the images of each camera and confirm the presence of a driver if they match. For example, if a person is visible in the upper half of the first camera's screen and also in the lower half of the second camera's screen, it could be considered that a driver is present. Alternatively, the logic could be that the positional relationship between the head guard in the first camera indicates the presence of a person, and the positional relationship between the steering wheel in the second camera indicates the presence of a person, in which case it could be determined that a driver is present.
[0327] In this way, the system can be equipped with a function to enable / disable the notification function depending on whether or not there is a driver, or by combining it with the "vehicle status" (engine ON / OFF, ignition ON / OFF, etc.) and the "presence or absence of people (workers) around the vehicle," the conditions for notification and the notification method can be finely configured. For example, if a driver is present, a first notification method (buzzer, etc.) will be used, and if not, a second notification method (police light, etc.) will be used to warn, or it is possible to configure the system to change the notification method considering two or more factors such as "presence or absence of a driver," "vehicle status," "presence or absence of people around the vehicle," and "location of people around the vehicle." Allowing users to arbitrarily set and change the enabling / disabling of the notification function and the notification method will enable flexible responses in actual operational settings.
[0328] In addition to the various image area processing methods described above ((2-1) to (2-5)), it is advisable to adopt a configuration in which the position of the imaging device 12 (Figure 4 (page 4)) is set by the user or by automatic detection, and then the presence or absence of a driver on the forklift (vehicle 400 in Figure 1 (page 1)) and the presence or absence of people in the surrounding area are detected and notified. This is expected to significantly improve work safety and operational efficiency. (2-7)(2-1) to (1-6) Furthermore, it is preferable to have a media recording function that records the output or internally generated information on the storage medium 500. Furthermore, the above-mentioned functions can be implemented as the execution (processing) of a program by the computer provided in the control unit 111. [Potential contribution to SDGs]
[0329] This invention relates to image processing technology centered on a camera system mounted on a vehicle (particularly a forklift, etc.), and detection and notification functions linked thereto. As shown below, it is believed that by utilizing these configurations, it is possible to contribute to multiple of the Sustainable Development Goals (SDGs) advocated by the United Nations. 1. Improve occupational safety and health (SDGs 3, 8) SDG 3: "Ensure healthy lives and promote well-being for all."
[0330] The image processing system of the present invention enables wide-area, real-time monitoring of the area surrounding work vehicles, including forklifts, and can provide advance warnings to drivers and surrounding workers of potential hazards. As a result, it contributes to reducing the risk of accidents and injuries, and improves people's health and well-being by enhancing the safety of the work environment. SDG 8: "Decent Work and Economic Growth"
[0331] Strengthening safety measures is essential for creating a workplace environment where workers can perform their duties with peace of mind. The advanced driver detection and surrounding object recognition provided by this invention will promote the prevention of industrial accidents, leading to increased productivity and a more fulfilling work environment. 2. Building sustainable industries and infrastructure (SDGs 9, 11) SDG 9: "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."
[0332] Acquiring spherical images using in-vehicle cameras and image conversion technologies that utilize overlapping areas enable advanced image analysis and AI integration, contributing to the promotion of DX (Digital Transformation) in industrial settings. Advanced safety management systems will boost the quality of industrial infrastructure, leading to a reduction in industrial accidents and improved operational efficiency. SDG 11: "Make cities and human settlements inclusive, safe, resilient and sustainable."
[0333] If the safety system of the present invention is introduced in warehouses, factories, or urban redevelopment areas where work vehicles are in operation, the risk of collisions and work-related accidents will be reduced, and a safer and more comfortable environment can be created for local residents and workers. 3. Efficient resource use and reduction of environmental impact (SDGs 12, 13) SDG 12: "Responsible Consumption and Production"
[0334] Improving the accuracy of driver detection and surrounding object detection will help reduce errors and excessive operation, thereby encouraging reductions in the operating time and fuel consumption of forklifts and other equipment. By reducing unnecessary energy consumption, resource use will be optimized, contributing to the realization of a circular economy. SDG 13: "Take concrete action to combat climate change."
[0335] Achieving energy-efficient operations contributes to reducing greenhouse gas emissions. The reduction of accident risk and improvement of operational efficiency according to this invention ultimately prevent waste of power sources (electricity and fuel), thus contributing to climate change countermeasures. 4. Safety education and awareness through innovation (SDGs 4, 17) SDG 4: "Quality Education for All"
[0336] By accumulating and analyzing the image data and detection results constructed using this invention, it is possible to create content that enhances the quality of safety education, such as educational materials similar to those used in dashcams and applications in VR simulations. It can also be used for accident prevention and safe driving training. SDG 17: "Partnerships for the Goals"
[0337] By having work vehicle manufacturers, system integrators, and AI vendors collaborate to establish a safety infrastructure and share expertise on a common platform, the standardization and widespread adoption of safety technologies can be promoted.
[0338] As described above, the image processing system and detection / notification technology of the present invention can address multiple SDGs goals by improving safety and efficiency related to work vehicles. It contributes to the realization of a sustainable society in various aspects, from improving the working environment to reducing environmental impact.
[0339] Furthermore, the scope of the present invention is not limited to the configurations explicitly described in the specification, but also includes combinations of various aspects of the present invention disclosed herein. While the configurations for which patent protection is sought are specified in the appended claims, we intend to include configurations disclosed herein that are not currently specified in the claims in the future.
[0340] The present invention is not limited to the configuration described in the embodiments described above. The components in each embodiment, example, and modification described above can be arbitrarily selected and combined. Furthermore, any component in each embodiment, example, and modification can be arbitrarily combined with any component described in the means for solving the invention, or any component that embodies any component described in the means for solving the invention. The present application intends to obtain rights to these as well through amendments or divisional applications. Even if there is a description such as "in the case of..." or "when...", it is not meant to be a configuration that is limited to that case or time. Configurations that do not fall under these cases or times are also disclosed, and the present application intends to obtain rights to them. Also, even if there is a sequence of descriptions, it is not limited to that order. Configurations with some parts deleted or the order rearranged are also disclosed, and the present application intends to obtain rights to them.
[0341] Furthermore, by converting to a design registration application, we intend to acquire rights to the overall design or a partial design. The drawing depicts the entire device with solid lines, but it is a drawing that includes not only the overall design but also partial designs claimed for parts of the device. For example, it is a drawing that includes not only a partial design for a part of the device's components, but also a partial design for a part of the device regardless of its components. A part of the device may be a component of the device, or a part of a component. We intend to acquire rights not only to the overall design, but also to a partial design where any part of the solid lines in the drawing is represented by dashed lines. In addition, all modules, components, and parts inside the device's casing that are shown in the drawing are independently tradable, and similarly, we intend to acquire rights to them by converting to a design registration application. [Explanation of Symbols]
[0342] 100...System, 11...Control device, 115...Image processing unit, 116...Detection unit, 12...Photography device, 420...Display unit, 430...Notification unit 30
Claims
1. An image processing system comprising image processing means for extracting one or more image regions such that a portion of an input image overlaps with the input image, and converting the extracted one or more image regions into an output image with a different region shape from the aforementioned image regions.
2. The image processing system according to claim 1, wherein the image processing means is configured to generate a display image by arranging a plurality of the output images.
3. The image processing system according to claim 1 or 2, wherein the image processing means is configured to extract the image region from an annular region of the input image, excluding the central circular region.
4. The image processing system according to claim 3, wherein the image processing means is configured to extract the image region by adjusting the ratio of the radial length in the annular region and the length of the arc located outside the radial direction, based on a preset aspect ratio in the display section of the displayed image.
5. The image processing system according to claim 1 or 2, wherein the image processing means is configured to accept the setting of the range of the overlapping portion in the image region based on a specification from the user.
6. The image processing system according to claim 1 or 2, wherein the image processing means is configured to change the range of the overlapping portion based on the attributes of the object to be detected, which have been set in advance.
7. The image processing system according to claim 1 or 2, wherein the image processing means is configured to change the range of the overlapping portion based on the attributes of the subject captured in the input image.
8. The image processing system according to claim 1 or 2, wherein the image processing means is configured to change the range of the overlapping portion based on the position of the subject captured in the input image.
9. The image processing system according to claim 1 or 2, wherein the image processing means is configured to change the range of the overlapping portion based on the shooting range of the input image.
10. The system further includes a detection means for detecting a pre-set target object from the input image, The image processing system according to claim 2, wherein the image processing means is configured to output information indicating the detected object in the region of the detected object in the displayed image when the detection means detects the detected object.
11. The system further includes a detection means for detecting a pre-set target object from the input image, The aforementioned input image was captured by an on-board camera mounted on the work vehicle. The image processing system according to claim 1 or 2, wherein the detection means is configured to detect the presence of the driver in the work vehicle by reading a predetermined worker code assigned to the driver of the work vehicle.
12. The image processing system according to claim 11, wherein the detection means is configured to change the evaluation region in the input image for detecting the presence of the driver based on the position of the on-board camera in the work vehicle.
13. The image processing system according to claim 11, wherein the detection means is configured to change an evaluation region in the input image for detecting the presence of the driver, based on its positional relationship with a specific structure fixed to and used on the work vehicle.
14. The image processing system according to claim 11, wherein the detection means is configured to perform detection processing on each of the multiple input images obtained from each of the multiple in-vehicle cameras, and to detect the presence of the driver based on the respective detection results.
15. The image processing system according to claim 11, wherein the detection means is configured to output the detection results when the detection result for the input image from the first in-vehicle camera and the detection result for the input image from the second in-vehicle camera are within a predetermined range.
16. The image processing system according to claim 11, wherein the detection means is configured to detect the presence of a driver when the relative position of the driver with respect to the surrounding structure detected with respect to the input image from the first in-vehicle camera and the relative position of the driver with respect to the surrounding structure detected with respect to the input image from the second in-vehicle camera are within a predetermined range.
17. The image processing system further includes a notification means, The image processing system according to claim 11, wherein the notification means is configured to notify the detection result of the object to be detected when the presence of the driver in the work vehicle is detected by the detection means.
18. The image processing system according to claim 17, wherein the notification means is configured to notify the detection result of the object to be detected in a different manner than when the presence of the driver is detected, when the presence of the driver is not detected by the detection means.
19. The image processing system further includes a notification means, The image processing system according to claim 11, wherein the notification means is configured to provide different notifications of the detection result of the detected object based on the positional relationship of the detected object with respect to the work vehicle when the detection means detects the detected object in the vicinity of the work vehicle.
20. The image processing system according to claim 19, wherein the notification means is configured to notify the detection result of the object in a first manner if the positional relationship between the object to be detected and the work vehicle is within a predetermined range arbitrarily set in advance, and to leave the detection result of the object to be detected unattended in a second manner if it is outside the predetermined range.
21. The image processing system further includes a notification means, The image processing system according to claim 11, wherein the notification means further comprises a configuration that causes the notification of the detection result of the object to be detected to vary based on at least one of the status of the work vehicle, the presence or absence of an object to be detected in the vicinity of the work vehicle, and the location of the object to be detected.
22. An image processing method comprising a configuration in which a computer extracts one or more image regions such that a portion of an input image overlaps, and performs an image processing step of converting the extracted one or more image regions into an output image with a different region shape from the aforementioned image regions.
23. A program for causing a computer to function as each of the means in the image processing system described in any one of claims 1 to 21.
Citation Information
Patent Citations
Imaging apparatus, system, and forklift equipped with the system
JP2017132298A