Image processing device, image processing method, and image processing program
The image processing device addresses the prolonged learning time issue by generating composite video data with non-overlapping object positions, efficiently reducing the number of synthesized video data to shorten training time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- OMRON CORP
- Filing Date
- 2024-11-07
- Publication Date
- 2026-05-19
AI Technical Summary
The prolonged learning time when using video data as teacher data for machine learning in monitoring systems is a challenge due to the inclusion of a time element indicating when abnormalities occur, which is not present when using still image data.
An image processing device generates composite video data by combining multiple object image data into a single background video data without overlapping positions, reducing the number of composite video data and training time through parameter narrowing and synthesis.
This approach suppresses the need for prolonged training using video data by generating composite video data efficiently, reducing the number of synthesized video data without compromising variations, thus shortening the learning time.
Smart Images

Figure 2026082534000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing method, and an image processing program.
Background Art
[0002] A monitoring system that monitors a road using video data captured on the road is being used. In Patent Document 1, a traffic control device using a learning model constructed by machine learning with video data of a road when an emergency occurred on the road in the past as input data has been proposed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When using still image data as teacher data, learning for the presence or absence of abnormalities may be performed for each still image data. On the other hand, when using video data as teacher data, a time element indicating in which time zone of the video data an abnormality has occurred will be included. Therefore, learning for the presence or absence of abnormalities will be performed after checking from the start to the end of the video data, and when using video data as teacher data, the time used for learning will be longer than when using still image data as teacher data. In machine learning, a large number of video data are used as teacher data, so the prolongation of learning becomes more prominent.
[0005] One aspect of the disclosed technology aims to provide an image processing apparatus, an image processing method, and an image processing program that can suppress the prolongation of learning using video data.
Means for Solving the Problems
[0006] One aspect of the disclosed technology is exemplified by the following image processing device. This image processing device includes a storage unit that stores background video data of a target area and object image data to be detected, and a control unit that receives the specification of parameters for a plurality of items, including the position in which the object image data is placed on the background video data, and generates composite video data by combining the object image data with the background video data according to each of the combinations of parameters for the plurality of items. For a first group of parameter combinations among the combinations of parameters for the plurality of items in which the positions in which the object image data is placed do not overlap, the control unit generates the composite video data by combining a plurality of object image data according to each of the first group of parameter combinations into a single background video data.
[0007] This image processing device generates composite video data by combining multiple object image data, according to each of the first set of parameter combinations, into a single background video data for the first set of parameter combinations where the positions of the object image data do not overlap. By generating composite video data in this way, the number of composite video data can be reduced without reducing the variations of object image data specified by the parameters. Therefore, this image processing device can suppress the prolonged learning time using composite video data.
[0008] The image processing device may have the following features: Each of the above multiple items includes multiple options, and the control unit uses an experimental design method corresponding to the number of items and the number of options among the combinations of parameters of the above multiple items to generate the above composite video data. The parameters for multiple items are narrowed down. An image processing device with this feature can reduce the number of synthesized video data generated by narrowing down the parameters for the multiple items used to generate the synthesized video data. In turn, it can suppress the length of time required for training using the synthesized video data.
[0009] This image processing device may have the following features: The above parameter includes a range of magnification ratios for the object image data, and the control unit performs a magnification process on the object image data using a magnification ratio selected from the above range of magnification ratios, and then synthesizes the object image data with the background video data. With such an image processing device, object image data of various sizes within the above magnification ratio range can be synthesized with the background video data. Therefore, the effort required of the user is reduced compared to specifying the size of the object image data individually.
[0010] This image processing device may have the following features: The above parameter includes a contrast adjustment range for the object image data, and the control unit performs contrast adjustment processing on the object image data using a contrast selected from the above contrast adjustment range, and then synthesizes the object image data with the background video data. With such an image processing device, object image data with various contrasts within the adjustment range can be synthesized with the background video data. Therefore, the effort required of the user is reduced compared to specifying the contrast of each object image data individually.
[0011] This image processing device may have the following features: The above parameter includes an adjustment range for the amount of rotation of the object image data, and the control unit rotates the object image data around a predetermined position by an amount selected from the adjustment range for the amount of rotation, and then synthesizes the resulting object image data with the background video data. With such an image processing device, object image data with various amounts of rotation within the adjustment range can be synthesized with the background video data. Therefore, the user's effort is reduced compared to specifying the amount of rotation for each object image data individually.
[0012] The technology being disclosed can also be understood from the perspective of image processing methods and image processing programs. [Effects of the Invention]
[0013] According to the disclosed technology, it is possible to suppress the need for prolonged training using video data. [Brief explanation of the drawing]
[0014] [Figure 1] Figure 1 shows an example of a falling object detection system according to an embodiment. [Figure 2] Figure 2 shows an example of the hardware configuration of the detection device according to this embodiment. [Figure 3] Figure 3 shows an example of a processing block of a detection device according to an embodiment. [Figure 4] Figure 4 shows an example of a background video model held by the management database. [Figure 5] Figure 5 shows an example of a background video model management table in the management database. [Figure 6] Figure 6 shows an example of a falling object image held in the management database. [Figure 7] Figure 7 shows an example of a fallen object management table in the management database. [Figure 8] Figure 8 shows an example of a scaling rate management table in the management database. [Figure 9] Figure 9 shows an example of a contrast management table in the management database. [Figure 10] Figure 10 shows an example of a location management table in the management database. [Figure 11] Figure 11 shows an example of a composite video pattern management table in the management database. [Figure 12] Figure 12 shows an example of the falling object parameter setting screen output by the reception unit in this embodiment. [Figure 13] Figure 13 shows an example of a composite video parameter setting screen output by the reception unit in this embodiment. [Figure 14]FIG. 14 is a diagram showing an example of a composite video pattern management table when placing falling object images at different positions in one composite video data. [Figure 15] FIG. 15 is a diagram showing an example of an orthogonal table generated by the creation unit in the embodiment. [Figure 16] FIG. 16 is a diagram showing an example of a composite video pattern management table with the number of composite video data reduced based on the experimental design method. [Figure 17] FIG. 17 is a diagram showing an example of composite video data generated by the generation unit in the embodiment. [Figure 18] FIG. 18 is a diagram showing an example of the processing flow of the detection device 2 according to the embodiment.
MODE FOR CARRYING OUT THE INVENTION
[0015] <APPLICATION EXAMPLE> The invention according to this application example is, for example, the detection device 2 illustrated in FIG. 1. The detection device 2 is, for example, a device that inputs video data captured by the camera 1 into a learning model 241 (see FIG. 3) to detect the presence or absence of a falling object. Further, the detection device 2 generates teacher data in the machine learning of the learning model 241 by synthesizing video data.
[0016] In the detection device 2, a background video model 271 (see FIG. 4) of the road 500 captured by the camera 1 is stored in advance in the auxiliary storage unit 203 (see FIG. 2). Further, a falling object image 273 of an object to be detected on the road 500 captured by the detection device 2 is also stored in advance in the auxiliary storage unit 203. The detection device 2 accepts the specification of parameters for a plurality of items including the position where the falling object image 273 is placed in the background video model 271. Then, the detection device generates composite video data 231 by synthesizing the falling object image 273 with the background video model 271 according to each combination of parameters of the plurality of items. The composite video data 231 is used, for example, as teacher data in the machine learning of the learning model 241.
[0017] In machine learning, it is preferable to prepare training data that reflects various conditions. However, if various conditions are reflected in the parameters, the number of parameter combinations becomes enormous. If one synthesized video data 231 is generated for each of these enormous parameter combinations, the number of synthesized video data 231 generated will also be enormous. If the learning model 241 is trained using an enormous number of synthesized video data 231, the training time will also be enormous.
[0018] Therefore, the detection device 2, for a first group of parameter combinations in which the positions where the falling object images 273 are placed do not overlap, synthesizes multiple falling object images 273 according to each of the first parameter combinations into a single background video model 271 to generate synthesized video data 231. By generating synthesized video data 231 in this way, the number of synthesized video data 231 can be reduced without reducing the variations of the falling object images 273 according to the specified parameters. Thus, according to this application example, the length of learning using video data can be suppressed.
[0019] <Embodiment> The embodiments will be described below with reference to the drawings. Figure 1 shows a falling object detection according to the embodiment. This figure shows an example of system 100. The falling object detection system 100 is a system that detects objects that have fallen onto a road 500. The falling object detection system 100 includes a camera 1, a detection device 2, and a display device 3. The camera 1 and the detection device 2 are connected by a computer network N1. The detection device 2 and the display device 3 are connected by a connection cable L1. Road 500 is, for example, an expressway or a national highway. Road 500 may also be a national highway, a prefectural road, a municipal road, or a private road.
[0020] Camera 1 is a video camera that films the road 500 to be monitored. Camera 1 may be a network camera that transmits the captured video data via a computer network N1, for example. Camera 1 is a digital video camera that employs a Charge Coupled Device (CCD) or Complementary Metal-Oxide-Semiconductor (CMOS) as its image sensor. The video data captured by Camera 1 is transmitted to the detection device 2 via the computer network N1. Road 500 is an example of a "target area".
[0021] The detection device 2 is an information processing device that detects fallen objects on the road 500 based on video data acquired from camera 1 via the computer network N1. The detection device 2 performs fallen object detection using, for example, a learning model constructed by machine learning. The learning model of the detection device 2 uses composite video data as training data, which is created by superimposing images of the fallen objects to be detected onto road video data previously captured by camera 1. For example, when the detection device 2 detects fallen objects on the road 500, it outputs the video data of the detected objects to the display device 3.
[0022] Display device 3 is a display monitored by traffic monitor K1. Display device 3 is, for example, a Liquid Crystal Display (LCD), Plasma Display Panel (PDP), inorganic electroluminescence (EL) panel, or organic EL panel. Display device 3 displays video data output by detection device 2.
[0023] Figure 2 shows an example of the hardware configuration of the detection device 2 according to the embodiment. The detection device 2 comprises a Central Processing Unit (CPU) 201, a main memory unit 202, an auxiliary memory unit 203, a communication unit 204, a connection terminal 205, and a bus B1. The CPU 201, the main memory unit 202, the auxiliary memory unit 203, the communication unit 204, and the connection terminal 205 are interconnected by the bus B1.
[0024] The CPU201 is also called a microprocessor unit (MPU) or processor. The CPU201 is not limited to a single processor and may be in a multiprocessor configuration. Furthermore, a single CPU201 connected via a single socket may have a multicore configuration. At least a portion of the processing performed by the CPU201 may be performed by other processors, such as dedicated processors like a Digital Signal Processor (DSP), Graphics Processing Unit (GPU), numerical processor, vector processor, or image processing processor. Also, at least a portion of the processing performed by the CPU201 may be performed by integrated circuits (ICs) or other digital circuits. Furthermore, at least a portion of the CPU201 may include analog circuits. Integrated circuits include Large Scale Integrated Circuits (LSIs), Application Specific Integrated Circuits (ASICs), and Programmable Logic Devices (PLDs). PLDs include, for example, Field-Programmable Gate Arrays (FPGAs). The CPU201 is a combination of a processor and integrated circuits. This is also acceptable. The combination is called, for example, a microcontroller unit (MCU), system-on-a-chip (SoC), system LSI, or chipset. In the detection device 2, the CPU 201 deploys the program stored in the auxiliary storage unit 203 to the work area of the main storage unit 202 and controls peripheral devices through program execution. This allows the detection device 2 to perform processing that matches a predetermined purpose. The main storage unit 202 and the auxiliary storage unit 203 are recording media that the detection device 2 can read. The CPU 201 is an example of a "control unit".
[0025] The main memory unit 202 is exemplified as a memory unit that is directly accessed by the CPU 201. The main memory unit 202 includes Random Access Memory (RAM) and Read Only Memory (ROM).
[0026] The auxiliary storage unit 203 stores various programs and data on a recording medium in a read-write manner. The auxiliary storage unit 203 is also called an external storage device. Multiple background video models may be stored in the auxiliary storage unit 203. A background video model is, for example, video data that shows only the background of a video of road 500 taken by camera 1. For example, a background video model is generated based on a video of road 500 taken by camera 1 when no vehicles are moving on it. A background video model is, for example, grayscale video data. In addition, multiple background video models taken under different shooting conditions are stored in the auxiliary storage unit 203. Shooting conditions include the time of shooting, the weather at the time of shooting, the season in which it was shot, etc. A background video model is an example of "background video data".
[0027] Furthermore, the auxiliary storage unit 203 may store images of falling objects that are to be detected by the falling object detection system 100. Examples of falling objects include cardboard boxes, tires, and plastic sheets. The falling object images are, for example, still image data showing the falling objects captured by the camera 1. The falling object images are, for example, grayscale still image data. In addition, for example, still images of each falling object may be prepared, taken with different sizes and contrasts.
[0028] Furthermore, the auxiliary storage unit 203 stores the operating system (OS), various programs, various tables, etc. The OS includes a communication interface program that handles data exchange with external devices connected via the communication unit 204. External devices include, for example, other information processing devices and external storage devices connected via a computer network. The auxiliary storage unit 203 may also be, for example, part of a cloud system, which is a group of computers on a network.
[0029] The auxiliary storage unit 203 is, for example, an Erasable Programmable ROM (EPROM), a Solid State Drive (SSD), a Hard Disk Drive (HDD), etc. Alternatively, the auxiliary storage unit 203 may be a Compact Disc (CD) drive, a Digital Versatile Disc (DVD) drive, a Blu-ray® Disc (BD) drive, etc. Furthermore, the auxiliary storage unit 203 may be provided by a Network Attached Storage (NAS) or Storage Area Network (SAN). The auxiliary storage unit 203 is an example of a "storage unit".
[0030] The communication unit 204 is, for example, an interface with the computer network N1. The communication unit 204 communicates with external devices such as the camera 1 via the computer network N1.
[0031] Connection terminal 205 is the connection terminal to which the connection cable L1 is connected. When the detection device 2 detects a falling object, it outputs an alarm to the display device 3 via the connection cable L1.
[0032] <Processing block of detection device 2> Figure 3 shows an example of a processing block of the detection device 2 according to an embodiment. The detection device 2 includes a reception unit 21, a creation unit 22, a generation unit 23, a learning unit 24, an acquisition unit 25, an extraction unit 26, and a management database (referred to as "management DB" in the figure) 27. The detection device 2 performs processing as each of its respective parts, such as the reception unit 21, creation unit 22, generation unit 23, learning unit 24, acquisition unit 25, extraction unit 26, and management database 27, by having the CPU 201 execute a computer program that has been expanded in executable form in the main memory unit 202.
[0033] The management database 27 is a database that stores various information related to the detection of fallen objects by the detection device 2. The management database 27 is built in, for example, the auxiliary storage unit 203. The management database 27 stores, for example, background video models of the road 500 without vehicles or fallen objects, captured by the camera 1 under various conditions. Figure 4 is a diagram showing an example of a background video model 271 held by the management database 27. As described above, the background video model 271 is video data of the road 500 without vehicles or fallen objects. The management database 27 manages multiple background video models 271 captured under various conditions (weather, season, time of day, etc.).
[0034] Figure 5 shows an example of a background video model management table 272 in the management database 27. The background video model management table 272 has the following items: "Background ID", "Condition", and "Path Name".
[0035] The "Conditions" field stores the shooting conditions under which the background video model 271 was captured. In the example in Figure 5, "Conditions" includes the fields "Season," "Time of Day," and "Weather." These fields store information indicating the season, time of day, and weather when the background image 11 was captured. The "Background ID" field stores an ID (Background ID) that uniquely identifies the background video model 271 captured by camera 1. The "Path Name" field stores, for example, the path name of the background video model 271 stored in the auxiliary storage unit 203. In the example in Figure 5, a relative path is stored in "Path Name," but an absolute path may also be stored in "Path Name." The background video model management table 272 associates the shooting conditions under which the background video model 271 was captured with the background video model 271 itself.
[0036] Furthermore, the management database 27 also manages images of fallen objects (also referred to as "fallen object images") taken by camera 1 or other imaging devices including cameras. Fallen object images are images taken by the imaging device of various fallen objects that may occur on the road 500. Examples of fallen objects that may occur on the road 500 include cardboard boxes, tires, and plastic sheets.
[0037] Figure 6 shows an example of a falling object image 273 held by the management database 27. In the example in Figure 6, a falling object image 273 is shown in which a cardboard box is photographed as the falling object. A falling object image 273 is, for example, an image obtained by removing the margins around the photographed falling object from an image taken of the falling object by a photographing device. As described above, the falling object image 273 is managed by the management database 27. A falling object image 273 is an example of "object image data".
[0038] Figure 7 shows an example of a fallen object management table 274 in the management database 27. The fallen object management table 274 has the following items: "Fallen Object ID", "Type", and "Path Name".
[0039] The "Falling Object ID" field stores an ID (Falling Object ID) that uniquely identifies the image of the falling object captured by camera 1. The "Type" field stores information indicating the type of falling object. In the example in Figure 7, three types of falling objects are listed: "cardboard box," "tire," and "plastic sheet." The "Path Name" field stores the path name of the falling object image 273 stored in the auxiliary storage unit 203. In the example in Figure 7, a relative path is stored in "Path Name," but an absolute path may also be stored in "Path Name."
[0040] Figure 8 shows an example of a scaling factor management table 275 contained in the management database 27. The scaling factor management table 275 has the following items: "Size ID", "Name", "MIN", and "MAX".
[0041] The "Size ID" stores an ID that uniquely represents the combination of "Name," "MIN," and "MAX" in the magnification management table 275. The "Name" stores the name corresponding to the combination of "MAX" and "MIN." "MIN" and "MAX" store information indicating the range of magnification for the falling object image 273. "MIN" stores information indicating the lower limit of the magnification range. "MAX" stores information indicating the upper limit of the magnification range. In the case of Figure 8, for example, if "Large Size" is specified, the falling object image 273 will be magnified in the range of 1.1 to 1.2 times.
[0042] Figure 9 shows an example of a contrast management table 276 contained in the management database 27. The contrast management table 276 has the following items: "Contrast ID", "Name", "MIN", and "MAX".
[0043] The "Contrast ID" stores an ID that uniquely represents the combination of "Name," "MIN," and "MAX" in the contrast management table 276. The "Name" stores the name corresponding to the combination of "MAX" and "MIN." "MIN" and "MAX" store information indicating the contrast adjustment range for the falling object image 273. "MIN" stores information indicating the lower limit of the width of the contrast adjustment range. "MAX" stores information indicating the upper limit of the contrast adjustment range. In the case of Figure 9, for example, if "Low Contrast" is specified, the contrast of the falling object image 273 is adjusted within a range of 1.1 to 1.2 times.
[0044] Figure 10 shows an example of a location management table 277 in the management database 27. The location management table 277 manages the locations of the falling object images 273 that are placed on the background video model 271. The location management table 277 has the following items: "Location ID", "Name", and "Location".
[0045] The "Location ID" contains an ID that uniquely identifies the combination of "Name" and "Location" in the location management table 277. The "Name" contains the name corresponding to the "Location". The "Location" contains the location corresponding to the near, medium, and far distances on the background video model 271.
[0046] Figure 11 shows an example of a composite video pattern management table 278 in the management database 27. The composite video pattern management table 278 manages various parameters of composite video data specified by the user. The composite video pattern management table 278 has the following items: "Pattern ID", "Video ID", "Position", "Type", "Size", "Contrast", and "Background Video Model".
[0047] The "Pattern ID" contains the parameters "Video ID", "Position", "Type", "Size", "Contrast", and "Background Video Model" from the composite video pattern management table 278. An ID that uniquely identifies the combination of elements is stored. The "Video ID" stores an ID that uniquely identifies the composite video data generated by the generation unit 23. The "Position" stores information indicating the position where the falling object image 273 is placed on the background video model 271. The information indicating the position is, for example, the position ID in the position management table 277. The "Type" stores information indicating the type of falling object image 273 placed on the background video model 271. The information indicating the type is, for example, the falling object ID in the falling object management table 274. The "Size" stores information indicating the size of the falling object image 273 placed on the background video model 271. The information indicating the size is, for example, the size ID in the magnification management table 275. The "Contrast" stores information indicating the contrast of the falling object image 273 placed on the background video model 271. The information indicating the contrast is, for example, the contrast ID in the contrast management table 276. The "Background Video" stores information indicating the background video model 271 used to generate the composite video data. The information indicating the background video model 271 is, for example, the background ID in the background video model management table 272. The combination of "position," "type," "size," "contrast," and "background video model" can be considered the placement conditions for positioning the falling object image 273. The composite video pattern management table 278 stores the combination of the falling object image 273 and the placement conditions for positioning the falling object image 273.
[0048] Returning to Figure 3, the reception unit 21, creation unit 22, generation unit 23, and learning unit 24 are processing blocks used, for example, to generate composite video data used for training the learning model 241 of the learning unit 24. The reception unit 21 accepts the setting of conditions for compositing the falling object image onto the background video model.
[0049] Figure 12 shows an example of a falling object parameter setting screen 211 output by the reception unit 21 in this embodiment. The reception unit 21 receives the specification of the magnification ratio and contrast of the falling object image 273 via the falling object parameter setting screen 211. The falling object parameter setting screen 211 includes a size specification area 2111 and a contrast specification area 2112. In the size specification area 2111, the ranges of magnification ratios for large, medium, and small sizes are specified. In the contrast specification area 2112, the adjustment ranges for high contrast, medium contrast, and low contrast are specified. The reception unit 21 stores the range of magnification ratios specified via the falling object parameter setting screen 211 in the magnification ratio management table 275. The reception unit 21 also stores the adjustment range of contrast specified via the falling object parameter setting screen 211 in the contrast management table 276.
[0050] Figure 13 shows an example of a composite video parameter setting screen 212 output by the reception unit 21 in this embodiment. The reception unit 21 accepts the selection of various parameters related to the generation of composite video data used for training the learning model 241. The composite video parameter setting screen 212 includes a falling object selection area 2121, a position selection area 2122, a size selection area 2123, a contrast selection area 2124, and a background specification area 2125. The falling object selection area 2121, the position selection area 2122, the size selection area 2123, the contrast selection area 2124, and the background specification area 2125 can be considered items of parameters to be set.
[0051] The falling object selection area 2121 accepts the selection of parameters indicating the type of falling object to be placed in the composite video data. In the example in Figure 13, the options selected in the falling object selection area 2121 include "cardboard box," "tire," and "plastic sheet." One or more falling object options can be specified in the falling object selection area 2121. In the example in Figure 13, "cardboard box," "tire," and "plastic sheet" are selected.
[0052] In the position selection area 2122, the selection of parameters indicating the position of falling objects to be placed in the composite video data is accepted. In the example in Figure 13, the position options include "short distance," "medium distance," and "long distance." In the position selection area 2122, one or more position options are accepted. The options are determined. In the example in Figure 13, all three options—"short distance," "medium distance," and "long distance"—are selected.
[0053] The size selection area 2123 accepts the selection of a parameter that indicates the size of the falling object to be placed in the composite video data. In the example in Figure 13, the size options are "Large," "Medium," and "Small." One or more size options can be specified in the size selection area 2123. In the example in Figure 13, all three sizes, "Large," "Medium," and "Small," are selected.
[0054] The contrast selection area 2124 accepts the selection of parameters that indicate the contrast of falling objects to be placed in the composite video data. In the example in Figure 13, the contrast options are "High," "Medium," and "Low." One or more contrast options can be specified in the contrast selection area 2124. In the example in Figure 13, all three options, "High," "Medium," and "Low," are selected.
[0055] In addition, while parameter selection is accepted via checkboxes in the falling object selection area 2121, position selection area 2122, size selection area 2123, and contrast selection area 2124, parameter selection may also be accepted by means other than checkboxes (e.g., radio buttons, text boxes, lists, etc.).
[0056] In the background specification area 2125, the selection of parameters indicating the background video model to be used to generate the composite video data is accepted. In the example in Figure 13, the background video model with the path name "~ / mov / mov1.mp4" is selected. The reception unit 21 stores the parameters selected in the composite video parameter setting screen 212 in the composite video pattern management table 278. At this stage, each "video ID" in the composite video pattern management table 278 is stored with a unique ID. As a result, the composite video pattern management table 278 becomes, for example, the state illustrated in Figure 11. The parameters selected by the falling object selection area 2121, position selection area 2122, size selection area 2123, and contrast selection area 2124 can be said to be parameters related to the falling object image 273.
[0057] The creation unit 22 creates combinations of parameters related to the generation of composite video data based on the composite video pattern management table 278 in which the parameters are stored by the reception unit 21. When parameters are selected on the composite video parameter setting screen 212 as exemplified in Figure 13, there are 3 types of falling objects, 3 types of positions, 3 types of sizes, and 3 types of contrasts, resulting in 81 possible parameter combinations. When composite video data is generated for each parameter combination, 81 composite video data files are generated.
[0058] The creation unit 22 reduces the number of synthesized video data to be generated by combining parameter combinations in which the positions of the falling object images 273 do not overlap into a single synthesized video data. In other words, if the falling object images 273 are in different positions, they will not overlap even if they are placed in a single synthesized video data. Therefore, even if a synthesized video data containing multiple falling object images 273 is used as training data in the machine learning of the learning model 241 described later, machine learning can be performed on each of the falling object images 273 placed in the synthesized video data. Accordingly, the creation unit 22 creates parameters related to the generation of synthesized video data so as to include falling object images 273 in different positions in a single synthesized video data.
[0059] Figure 14 shows an example of a composite video pattern management table 278 when falling object images 273 at different locations are placed in a single composite video data. By placing falling object images 273 at three different locations (near, medium, and far) into a single composite video data, the number of generated composite video data is reduced from 81 to 27. In Figure 14, the parameter group associated with the common video ID is "the combination of the first parameter". This is an example of a "group of first parameters". For example, the parameters of pattern IDs "1", "28", and "55" associated with video ID "1" are an example of a "group of first parameter combinations".
[0060] The creation unit 22 may further reduce the number of synthesized video data to be generated based on the experimental design method. In this embodiment, the parameters of the falling object image 273 placed in the synthesized video data are four factors: position, type, size, and contrast. The position factor has three levels: near, medium, and far. The type factor has three levels: cardboard box, tire, and vinyl sheet. The size factor has three levels: large, medium, and small. The contrast factor has three levels: high, medium, and low. In other words, in this embodiment, there are three levels for each factor.
[0061] Therefore, the creation unit 22 reduces the number of synthesized video data based on the four-factor, three-level parameters specified by the synthesized video parameter setting screen 212, based on the experimental design method. The creation unit 22 may also reduce the number of synthesized video data based on, for example, a four-factor, three-level orthogonal array.
[0062] Figure 15 shows an example of an orthogonal array 221 generated by the creation unit 22 in this embodiment. The orthogonal array 221 is an orthogonal array with 4 factors and 3 levels. In this embodiment, it is assumed that there are no interactions between the four factors: position, type, size, and contrast. Therefore, the creation unit 22 creates the orthogonal array 221 assuming that there are no interactions.
[0063] In this embodiment, as described above, multiple falling objects at different locations are placed in a single composite video data. Therefore, the creation unit 22, based on the orthogonal array 221 created based on the experimental design method, places the falling objects at three different locations into a single composite video data, further reducing the number of composite video data.
[0064] Figure 16 shows an example of a composite video pattern management table 278 in which the number of composite video data has been reduced based on experimental design. The creation unit 22 reduces the number of composite video data generated from 27 to 9 by reducing the number of composite video data using experimental design.
[0065] Returning to Figure 3, the generation unit 23 generates composite video data for each selected background video model 271 according to the parameters related to the falling object image 273, based on the composite video pattern management table 278 generated by the creation unit 22. The generation unit 23 generates composite video data for each video ID stored in the composite video pattern management table 278. Here, the generation of composite video data will be explained using the generation of composite video data for video ID "1" as an example.
[0066] The generation unit 23 refers to the composite video pattern management table 278 and obtains the background video model ID "B1" corresponding to video ID "1". The generation unit 23 then obtains the background video model 271 with the path name corresponding to background video model ID "B1" from the background video model management table 272.
[0067] The generation unit 23 refers to the composite video pattern management table 278 to obtain the type of falling object "901", size "D1", and contrast "C1" to be placed at position ID "P1". The generation unit 23 refers to the position management table 277 to obtain the position corresponding to position "P1". The generation unit 23 refers to the falling object management table 274 to obtain the falling object image 273 with the path name corresponding to the falling object type "901". In addition, the generation unit 23 obtains the range of magnification corresponding to size "D1" from the magnification management table 275, and obtains the range of contrast corresponding to contrast "C1" from the contrast management table 276.
[0068] The generation unit 23 generates an enlarged image of the falling object 273 obtained from the magnification management table 275. The image is enlarged or reduced at any magnification within a specified range. The generation unit 23 then performs contrast adjustment on the background video model 271 using a contrast value within a specified range obtained from the contrast management table 276. Here, "any magnification" is, for example, randomly determined within the range of acquired magnifications. The same applies to contrast.
[0069] The generation unit 23 then places the enlarged / reduced and contrast-adjusted falling object images 273 at position "P1" on all frame images of the background video model 271. The same process is followed for the falling object images 273 to be placed at positions "P2" and "P3," respectively.
[0070] Figure 17 shows an example of composite video data 231 generated by the generation unit 23 in an embodiment. In the composite video data 231 illustrated in Figure 17, a falling object O1 located at close range, a falling object O2 located at medium range, and a falling object O3 located at far range are placed on the background video model 271. Each of the falling objects O1, O2, and O3 is a falling object image 273 that has been enlarged / reduced and contrast adjusted by the generation unit 23 and placed on the background video model 271.
[0071] The composite video data 231 illustrated in Figure 17 corresponds, for example, to video ID "1" in the composite video pattern management table 278 illustrated in Figure 16. In this case, the falling object O1 placed in the foreground (close distance) has position "P1", type "901", size "D1", and contrast "C1". The falling object O3 placed in the farthest distance has position "P3", type "901", size "D3", and contrast "C2". The falling object O2 placed between falling objects O1 and O3 has position "P2", type "901", size "D2", and contrast "C3". The generated composite video data 231 is used, for example, as training data in machine learning for the learning model 241.
[0072] Returning to Figure 3, the learning unit 24 uses the synthesized video data 231 generated by the generation unit 23 as training data to perform machine learning on the learning model 241. In this embodiment, the falling object images 273 are placed at three different locations (positions P1, P2, and P3) of the background video model 271 to generate the synthesized video data 231. Therefore, the learning unit 24 can have the learning model 241 learn about detecting falling objects at three different locations using a single synthesized video data 231.
[0073] The acquisition unit 25 acquires video data of the road 500 captured by camera 1 from camera 1. The extraction unit 26 inputs the video data acquired by the acquisition unit 25 into the learning model 241 to detect fallen objects in the video data. The extraction unit 26 outputs the detection result to the display device 3, for example.
[0074] <Processing Flow> Figure 18 is a diagram showing an example of the processing flow of the detection device 2 according to the embodiment. Figure 18 illustrates the processing flow related to the generation of synthesized video data 231 by the detection device 2. Hereinafter, an example of the processing flow of the detection device 2 will be described with reference to Figure 18.
[0075] In step S1, the reception unit 21 outputs a falling object parameter setting screen 211 and accepts parameters related to the magnification and contrast. The falling object parameter setting screen 211 then outputs a composite video parameter setting screen 212 and accepts the specification of parameters related to the type of falling object, the position and size of the falling object, contrast, and background video model.
[0076] In step S2, the creation unit 22 stores the parameters received in step S1 in the composite video pattern management table 278.
[0077] In step S3, the creation unit 22 updates the composite video pattern management table 278 by reducing the number of parameter combinations stored in the composite video pattern management table 278 in step S2 by placing falling objects whose positions do not overlap into a single composite video data. Furthermore, the creation unit 22 may update the composite video pattern management table 278 by further reducing the number of parameter combinations stored in the composite video pattern management table 278 based on the experimental design method.
[0078] In step S4, the generation unit 23 generates synthetic video data 231 by referring to the synthetic video pattern management table 278, which has had its parameter combinations reduced in step S3. The generated synthetic video data 231 is used, for example, as training data for machine learning of the learning model 241.
[0079] <Effects of the Embodiment> According to this embodiment, the number of falling object images 273 whose positions do not overlap is placed in a single composite video data, thereby reducing the number of combinations of parameters selected by the composite video parameter setting screen 212. As a result, the number of generated composite video data 231 is reduced. By reducing the number of composite video data 231, the machine learning process of the learning model 241 is made more time-consuming.
[0080] In this embodiment, the combination of parameters selected by the synthesized video parameter setting screen 212 is further reduced by using experimental design. Therefore, according to this embodiment, the time required for machine learning of the learning model 241 is further suppressed.
[0081] In this embodiment, the magnification range is stored in the magnification management table 275 for each of "large," "medium," and "small." When one of "large," "medium," or "small" is selected, an arbitrary (random) magnification within the range stored in the magnification management table 275 is applied to the falling object image 273, and then it is composited onto the background video model 271. Therefore, according to this embodiment, the effort required from the user is reduced compared to specifying the size of the falling object image 273 individually.
[0082] In this embodiment, the contrast adjustment ranges for "high contrast," "medium contrast," and "low contrast" are stored in the contrast management table 276. When one of "high contrast," "medium contrast," or "low contrast" is selected, an arbitrary (random) contrast within the adjustment range stored in the contrast management table 276 is applied to the falling object image 273, and then it is composited with the background video model 271. Therefore, according to this embodiment, the effort required from the user is reduced compared to specifying the contrast of each falling object image 273 individually.
[0083] <Variation> In the embodiment described above, the falling object image 273 was placed in all frames of the background video model 271 according to the parameters relating to the falling object image 273. Here, a time element may be added to the parameters relating to the falling object image 273. That is, a time item may be added to the composite video pattern management table 278.
[0084] In such cases, the falling object images 273 may be arranged according to the first parameter for the first group of frame images belonging to the first period among the frame images of the background video model 271, and according to the second parameter for the second group of frame images belonging to the second period among the frame images of the background video model 271. By considering the element of time in this way, for example, the first half and second half of the background video model 271 may be different. The falling object image 273, for which parameters have been selected, can be placed. For example, by superimposing the falling object image 273 according to the parameter combination of video ID "1" in Figure 16 onto the first half of the playback time of the background video model 271, and superimposing the falling object image 273 according to the parameter combination of video ID "2" in Figure 16 onto the latter half of the background video model 271, the number of composite video data 231 can be further reduced.
[0085] Furthermore, in the embodiments described above, the falling object image 273 was placed in all frames of the background video model 271 according to the parameters relating to the falling object image 273. However, some frames of the background video model 271 may include normal frames (i.e., frames in which the falling object image 273 is not placed).
[0086] Furthermore, in the embodiments described above, parameter ranges were specified for the magnification ("large," "medium," "small") and contrast ("high," "medium," "low"). However, instead of specifying ranges, a single numerical value may be specified. In other words, the generation unit 23 may apply a specified numerical value rather than randomly determining the magnification and contrast from the specified ranges.
[0087] In the embodiment described above, the device for detecting objects that have fallen onto the road 500 and the device for generating the composite video data 231 are realized by a single detection device 2. However, the device for detecting objects that have fallen onto the road 500 and the device for generating the composite video data 231 may be separate devices. That is, there may be a detection device for detecting objects that have fallen onto the road 500 and a generation device for generating the composite video data 231.
[0088] In the embodiments described above, the parameters set for the falling object image 273 were specified as magnification and contrast. However, the parameters set for the falling object image 273 are not limited to magnification and contrast. For example, a rotation amount may be set for the falling object image 273. Similar to the magnification and contrast, an adjustment range may be specified for the rotation amount. The rotation direction may also be specified by the sign of the rotation amount. For example, a positive rotation amount may specify clockwise rotation, and a negative rotation amount may specify counterclockwise rotation. The creation unit 22 then rotates the falling object image 273 around a predetermined point (e.g., the center point) by a rotation amount selected from the adjustment range (e.g., +20 degrees). The creation unit 22 then generates composite video data 231 in which the rotated falling object image 273 is placed on the background video model 271 with the specified magnification and contrast.
[0089] On road 500, falling objects can occur in various orientations. By specifying the rotation amount as a parameter, composite video data 231 can be generated in which images 273 of falling objects in various orientations are placed on a background video model 271. Furthermore, within the above adjustment range, images 273 of falling objects with various rotation amounts can be composited onto the background video model 271. Therefore, the user's effort is reduced compared to specifying the rotation amount of each falling object image 273 individually.
[0090] The order in which the magnification, contrast, and rotation are applied to the falling object image 273 can be any order. Furthermore, some of the parameters among the magnification, contrast, and rotation may be applied to the falling object image 273.
[0091] In the embodiments described above, the falling object detection system 100 was described as an example, but the systems to which the disclosed technology is applied are not limited to the falling object detection system 100. The disclosed technology can be applied to various systems that detect anomalies by image processing, for example. Examples of systems to which the disclosed technology can be applied include intrusion detection systems that detect intrusion by people or animals, and monitoring systems that monitor equipment deterioration.
[0092] The embodiments and variations disclosed above can be combined in any way.
[0093] <Computer-readable recording medium> An information processing program that enables a computer or other machine or device (hereinafter referred to as "computer, etc.") to perform any of the above functions can be recorded on a recording medium that the computer, etc. can read. By having the computer, etc. read and execute the program on this recording medium, it can be made to provide that function.
[0094] Here, a recording medium that can be read by a computer refers to a recording medium that stores information such as data and programs through electrical, magnetic, optical, mechanical, or chemical means and can be read by a computer. Examples of such recording media that can be removed from a computer include flexible disks, magneto-optical disks, Compact Disc Read Only Memory (CD-ROM), Compact Disc-Recordable (CD-R), Compact Disc-ReWriterable (CD-RW), Digital Versatile Disc (DVD), Blu-ray Disc (BD), Digital Audio Tape (DAT), 8mm tape, flash memory, external hard disk drives, and Solid State Drives (SSDs). In addition, recording media that are fixed to a computer include internal hard disk drives, SSDs, and ROMs.
[0095] <Note 1> A storage unit (203) that stores background video data (271) of the target area (500) and image data (273) of the object to be detected, The system includes a control unit (201) that receives the specification of parameters for a plurality of items, including the position for placing the object image data (273) in the background video data (271), and generates composite video data (231) by combining the object image data (273) with the background video data (271) according to each of the combinations of parameters for the plurality of items, The control unit (201) is, For the first group of parameter combinations among the multiple parameter combinations in which the positions for placing the object image data (273) do not overlap, the multiple object image data (273) according to each of the first parameter combinations are combined into a single background video data (271) to generate the combined video data (231). Image processing device (2). <Note 2> Each of the aforementioned items includes multiple options. The control unit (201) narrows down the parameters of the multiple items used to generate the synthesized video data (231) from the combination of parameters of the multiple items by an experimental design method corresponding to the number of items and the number of choices. The image processing device (2) described in Appendix 1. <Note 3> The parameter includes a range of magnification for the object image data (273), The control unit (201) performs an enlargement process on the object image data (273) at a magnification selected from the range of magnifications, and then synthesizes the object image data (273) with the background video data (271). The image processing device (2) described in Appendix 1. <Note 4> The parameter includes the adjustment range for the contrast of the object image data (273), The control unit (201) selects the contrast from the adjustment range of the contrast. After performing contrast adjustment processing on the object image data (273), the object image data (273) is combined with the background video data (271). An image processing device as described in any one of the appendices 1 to 3. <Note 5> The parameter includes an adjustment range for the amount of rotation of the object image data (273), The control unit (201) combines the object image data (273), which has been rotated by an amount selected from the adjustment range of the rotation amount, around a predetermined position of the object image data (273), with the background video data (271). An image processing apparatus as described in any one of the appendices 1 to 4. <Note 6> A computer comprising a storage unit (203) that stores background video data (271) of a target area (500) and object image data (273) to be detected, which accepts the specification of parameters for a plurality of items, including the position in which the object image data (273) is placed on the background video data (271), and generates composite video data (231) by combining the object image data (273) with the background video data (271) according to each of the combinations of parameters for the plurality of items, For the first group of parameter combinations among the multiple parameter combinations in which the positions for placing the object image data (273) do not overlap, the multiple object image data (273) according to each of the first parameter combinations are combined into a single background video data (271) to generate the combined video data (231). Image processing methods. <Note 7> A computer comprising a storage unit (203) that stores background video data (271) of a target area (500) and object image data (273) to be detected, which accepts the specification of parameters for multiple items, including the position in which the object image data (273) is placed on the background video data (271), and generates composite video data (231) by combining the object image data (273) with the background video data (271) according to each combination of the parameters for the multiple items, Among the combinations of parameters of the aforementioned multiple items, for the first group of parameter combinations in which the positions for placing the object image data (273) do not overlap, the multiple object image data (273) according to each of the first parameter combinations are combined into a single background video data (271) to generate the combined video data (231). Image processing program. [Explanation of symbols]
[0096] 1. Camera 2. Detection device 3...Display device 21. Reception Department 22. Creation Department 23...Generation part 24. Learning Department 25...Acquisition part 26...Extraction part 27. Management Database 100. Falling Object Detection System 201··CPU 202...Main memory 203...Auxiliary storage section 204 Communications Department 205...Connection terminals 211. Falling Object Parameter Setting Screen 212. Synthetic video parameter settings screen 221··Orthogonal array 231 ··Composite video data 241 ··Learning Model 271 ··Background video model 272 ··Background video model management table 273 ··Image of falling object 274 ·· Falling Object Management Table 275 Magnification Management Table 276. Contrast Management Table 277. Location Management Table 278 ··Composite Video Pattern Management Table 500...road 2111 ··Size specification area 2112 ··Contrast specification area 2121 ··Falling object selection area 2122 ··Position selection area 2123 ··Size selection area 2124 ··Contrast selection area 2125...Background specification area B1...Bus K1...Traffic warden L1 Connection Cable N1 Computer Network O1 ··Falling object O2 ··Falling object O3 ··Falling object
Claims
1. A storage unit that stores background video data of the target area and image data of the object to be detected, The system includes a control unit that receives the specification of parameters for a plurality of items, including the position for placing the object image data in the background video data, and generates composite video data by combining the object image data with the background video data according to each combination of the parameters for the plurality of items, The control unit, Among the combinations of parameters of the aforementioned multiple items, for the first group of parameter combinations in which the positions for which the object image data is placed do not overlap, the multiple object image data according to each of the first parameter combinations are combined into a single background video data to generate the composite video data. Image processing device.
2. Each of the aforementioned items includes multiple options. The control unit narrows down the parameters of the multiple items to be used for generating the synthesized video data, using an experimental design method that corresponds to the number of items and the number of choices among the combinations of parameters of the multiple items. The image processing apparatus according to claim 1.
3. The aforementioned parameter includes the range of the magnification of the object image data, The control unit performs an enlargement process on the object image data using a magnification selected from the range of magnifications, and then synthesizes the object image data with the background video data. The image processing apparatus according to claim 1.
4. The parameter includes the adjustment range for the contrast of the object image data. The control unit performs contrast adjustment processing on the object image data using a contrast selected from the contrast adjustment range, and then synthesizes the object image data with the background video data. The image processing apparatus according to any one of claims 1 to 3.
5. The aforementioned parameter includes an adjustment range for the amount of rotation of the object image data. The control unit synthesizes the object image data, which has been rotated by an amount selected from the adjustment range of the rotation amount, around a predetermined position of the object image data, with the background video data. The image processing apparatus according to any one of claims 1 to 4.
6. A computer comprising a storage unit that stores background video data of a target area and image data of an object to be detected, which accepts the specification of parameters for a plurality of items, including the position in which the object image data is placed on the background video data, and generates composite video data by combining the object image data with the background video data according to each combination of the parameters for the plurality of items, Among the combinations of parameters of the aforementioned multiple items, for the first group of parameter combinations in which the positions for which the object image data is placed do not overlap, the multiple object image data according to each of the first parameter combinations are combined into a single background video data to generate the composite video data. Image processing methods.
7. A memory device that stores background video data of the target area and image data of the object to be detected. A computer having a memory section, which accepts the specification of parameters for a plurality of items, including the position for placing the object image data in the background video data, and generates composite video data by combining the object image data with the background video data according to each of the combinations of parameters for the plurality of items, Among the combinations of parameters of the aforementioned multiple items, for the first group of parameter combinations in which the positions for which the object image data is placed do not overlap, the multiple object image data according to each of the first parameter combinations are combined into a single background video data to generate the composite video data. Image processing program.