Image generation device, video analysis system, image generation method and program
The image generation device addresses neural network accuracy issues by generating training images based on user-specific conditions, enabling early and accurate video analysis without requiring post-installation data.
Patent Information
- Application Number
- JP2023161598
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-09-25
AI Technical Summary
Neural networks for video analysis may not achieve sufficient accuracy due to differences in input video trends between training and operation, particularly influenced by shooting environment and statistical trends in analysis targets, necessitating additional learning with user-specific data.
An image generation device that acquires use condition-related information to generate images expected during operation, using a video analysis system to create description data and images for training models, allowing early tuning without relying on post-installation user data.
Enables early and accurate video analysis by generating training images aligned with user conditions, ensuring high analysis accuracy from the start of device operation.
Smart Images

Figure 00000015_0000 
Figure 00000015_0001 
Figure 00000016_0000
Abstract
Description
[Technical Field]
[0001] The present invention relates to machine learning. [Background technology]
[0002] In recent years, the application of technology for analyzing surveillance camera footage has progressed. For example, technology for estimating a person's entire body posture has been proposed, and applications are progressing in the fields of customer safety in stores and urban surveillance. Machine learning, particularly learning models such as neural networks, is widely used in video analysis technology. Patent Document 1 describes a method for fine-tuning a trained neural network model using test data (video and labels) acquired by a user-side device. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 142766 Summary of the Invention [Problem to be solved by the invention]
[0004] Neural networks for video analysis may not be able to achieve sufficient analysis accuracy if the input video trends differ between when they are trained (developed) and when they are used by a user (operated). Examples of input video trends include those related to the shooting environment, such as lighting conditions, and statistical trends in the attributes of the analysis target, such as people in their 50s (e.g., many people are in their 50s). The method using test data, as described in Patent Document 1, requires additional learning using video captured over a certain period of time at the user's installation location. In other words, there is a problem in that sufficient analysis accuracy cannot be achieved during the test data creation period.
[0005] In view of the above problems, the present invention aims to enable tuning of a camera to begin earlier than a method in which additional learning is performed based on test data obtained after the camera is installed in a specific location. [Means for solving the problem]
[0006] The image generating device of the present invention includes: an acquisition means for acquiring, based on the use conditions of video analysis, related information related to the use conditions; Item information corresponding to a predetermined description item is obtained from the related information, and the item information Based on this, when the video analysis is used, an image that is expected to be the subject of the video analysis is displayed. and described according to the function of the video analysis. The video analysis system is characterized by comprising a description generating means for generating description data, and an image generating means for generating an analysis image to be used in the video analysis based on the description data. [Effects of the Invention]
[0007] According to the present invention, tuning of the camera can be started earlier than with a method in which additional learning is performed based on test data obtained after the camera is installed in a specific location. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of a video analysis system. [Figure 2] FIG. 2 illustrates an example of a hardware configuration of a learning device. [Figure 3] FIG. 2 is a diagram illustrating an example of a functional configuration of a learning device. [Figure 4] 10 is a flowchart illustrating a process executed by the learning device. [Figure 5] 10 is a flowchart showing details of an image description data generation process. [Figure 6] FIG. 10 is a diagram illustrating an example of order information. [Figure 7] FIG. 10 is a diagram illustrating an example of a table in which sub-items are defined. [Figure 8] FIG. 2 is a diagram illustrating an example of a functional configuration of a learning device. [Figure 9]10 is a flowchart showing details of an image description data correction process. [Figure 10] FIG. 10 is a diagram illustrating an example of a feedback screen. [Figure 11] FIG. 10 is a diagram showing an example of an image used for learning. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, each embodiment will be described with reference to the accompanying drawings.
[0010] <Embodiment 1> FIG. 1 shows an example of a video analysis system according to this embodiment. The video may be a video or a still image. The following mainly describes an example in which the video is a video. Video analysis is a process in which a subject to be analyzed is extracted from information captured in the video, and the movement and position of the subject are identified, thereby outputting an analysis result according to the video analysis function. The following mainly describes an example in which the subject is a person. In this embodiment, video analysis is performed by inputting the video into a video analysis model.
[0011] As shown in Fig. 1, the video analysis system includes various devices provided in a user environment 101 of a user who uses a video analysis device 104, and a learning device 111 provided in a vendor environment 110 of a developer who develops a video analysis model. The various devices provided in the user environment 101 and the learning device 111 are connected to the Internet 112. Various websites provided in a web environment 105 and a search server 109 are connected to the Internet 112. The learning device 111 transmits and receives data to and from the various devices connected to the Internet 112.
[0012] The video analysis device 104 is an integrated device that combines a camera (e.g., a surveillance camera) that captures video and a device that inputs the captured video into a video analysis model to perform video analysis. The video analysis device 104 also has functions according to the purpose of the video analysis. The learning device 111 optimizes the video analysis model to match these functions. In this embodiment, the video analysis device 104 has a fall detection function, and when the video analysis device 104 is in operation, the fall detection function is used to detect a person's fall from the captured video (input video). Note that the video analysis device 104 may also have a purchase analysis function, an intrusion detection function, a crowd detection function, a suspicious behavior detection function, etc., instead of or in addition to the fall detection function. Furthermore, although the present embodiment describes the video analysis device 104 being used in a store, it may also be used outdoors, such as in buildings other than stores or in town.
[0013] In this embodiment, the learning device 111 acquires information related to the use conditions (use condition related information) from the use conditions of the video analysis device 104, and generates an image (analysis image) representing an input video expected when the video analysis device 104 is in operation based on the acquired use condition related information. The learning device 111 then uses the image (analysis image) as a learning image for a video analysis model. In other words, before the video analysis device 104 begins operation, it is possible to prepare learning images that are close to those that will be used when the video analysis device 104 is in operation. The learning device 111 is an example of an image generation device.
[0014] The user environment 101 includes an internal mail server 102, a sales data management server 103, a video analysis device 104, and a PC 115 used by the user. The internal mail server 102 and the sales data management server 103 store various data related to the user's business and operations. The PC (personal computer) 115 is a terminal device used by the user. The PC 115 displays a UI screen received from the learning device 111 on a display device (not shown) and transmits information entered on the UI screen via an input device (not shown) to the learning device 111. The Web environment 105 includes user sites 106 established by users themselves, news sites 107 established by news organizations, and information sites 108 that compile various information. The search server 109 searches various websites provided in the Web environment 105 using the query (keywords) received from the learning device 111, and transmits the search results (URLs, etc.) to the learning device 111.
[0015] 2 is a block diagram showing the hardware configuration of the learning device 111. The learning device 111 has a CPU 201, a ROM 202, a RAM 203, a secondary storage device 204, an input device 205, a display device 206, and a network I / F 207. These components are connected via a bus 208 to input and output data to and from each other.
[0016] The CPU (Central Processing Unit) 201 controls the entire learning device 111. A GPU (Graphics Processing Unit) may be used instead of or together with the CPU. The ROM 202 is a non-volatile memory that stores various programs and data. The RAM 203 is a volatile memory that stores frame image data and temporary data generated in the flowcharts described below. The secondary storage device 204 is a rewritable storage device such as a hard disk drive or flash memory that stores image information, programs, various setting information, and the like.
[0017] The input device 205 is a keyboard, mouse, or the like, and is used to input operations from the developer. The display device 206 is a liquid crystal display, or the like, and displays processing results, etc., to the developer. The network I / F 207 is a modem, LAN, or the like that connects to a network such as the Internet 112 or an intranet. The CPU 201 reads out programs stored in the ROM 202, secondary storage device 204, or the like into the RAM 203 and executes them, thereby realizing processing corresponding to each step of the flowchart described below.
[0018] 3 is a block diagram showing the functional configuration of the learning device 111 according to this embodiment. The CPU 201 reads out a program stored in the ROM 202, the secondary storage device 204, etc. into the RAM 203 and executes it, thereby realizing each function shown in FIG.
[0019] The usage condition input unit 301 inputs the usage conditions of the video analysis device 104 used by the user. The related information acquisition unit 302 collects information related to the use conditions input in the use condition input unit 301 (use condition related information) from the user environment 101 and the Web environment 105 via the network I / F 207 . The related information storage unit 303 stores the use condition related information collected by the related information acquisition unit 302 in the secondary storage device 204, the RAM 203, or the like. The image description generating unit 304 generates image description data based on the use condition related information stored in the related information storage unit 303 . The image description storage unit 305 stores the image description data generated by the image description generation unit 304 in the secondary storage device 204, the RAM 203, or the like.
[0020] The image generation unit 306 generates an image and a correct label to be assigned to the image based on the image description data stored in the image description storage unit 305 . The generated image storage unit 307 stores the generated image generated by the image generation unit 306 and the correct label in the secondary storage device 204, the RAM 203, or the like. The trained model storage unit 308 stores trained video analysis models corresponding to the functions (here, the fall detection function) of the video analysis device 104 in the secondary storage device 204, etc. Note that known machine learning techniques such as neural networks are used as the video analysis models. The additional learning unit 309 performs additional learning of the video analysis model stored in the trained model storage unit 308 using the generated images and correct labels stored in the generated image storage unit 307. The model providing unit 310 transmits the video analysis model obtained by additional learning to the video analysis device 104 via the network I / F 207.
[0021] 4 is a flowchart showing the details of the processing executed by the learning device 111 according to this embodiment. In the following description, each process (step) is represented by adding an S to the beginning, and the notation of the process (step) is omitted.
[0022] First, in S401, the usage condition input unit 301 acquires order information for the video analysis device 104. Then, the CPU 201 extracts the usage conditions of the video analysis device 104 to be used by the user from the acquired order information, and acquires the extracted usage conditions.
[0023] FIG. 6 shows an example of order information when the video analysis device 104 is used in a store. The order information is generated when the developer receives an order for the video analysis device 104 from a user, and is stored, for example, in a data server installed in the vendor environment 110. The usage conditions input unit 301 acquires the order information from the data server via the network I / F 207. The order information in FIG. 6 includes information corresponding to various items such as an order recipient 601, a function 602, a business type 603, a store used 604, a camera installation location 605, a camera used 606, camera settings such as the set resolution 607, operating hours 608, a price 609, and a delivery date 610. The order recipient 601 indicates the name of the store where the video analysis device 104 will be used. The function 602 indicates the video analysis function corresponding to the intended use of the video analysis device 104. The business type 603 indicates the type of the store where the video analysis device 104 will be used. The store used 604 indicates the name of the branch of the store where the video analysis device 104 will be used. The camera used 606 indicates the model information of the camera used to capture the video.
[0024] In the following description, it is assumed that content corresponding to predetermined items related to the use conditions (items 601 to 610 in the example of FIG. 6) is extracted as the use conditions from among the items of the order information. Note that content corresponding to only some of the items 601 to 610 may be extracted. Furthermore, instead of the method of extracting the use conditions from the order information, the use condition input unit 301 may transmit a UI screen (not shown) for inputting the use conditions to the PC 115, and obtain information input by the user via the UI screen as the use conditions.
[0025] In S402, the related information acquisition unit 302 collects use condition related information based on the use conditions acquired in S401. The use condition related information includes the conditions of the image that is expected to be captured when shooting under the use conditions acquired in S401, and information for identifying those conditions. The image description storage unit 305 stores the collected use condition related information in the secondary storage device 204, RAM 203, etc. as needed.
[0026] The usage condition-related information relates to the background of the image. For example, it relates to background objects, lighting conditions, shooting range, and image quality. Background objects are the types of objects expected to exist in the background, such as product shelves in a store or roads in a city. Furthermore, lighting conditions are information specifying the sunlight conditions (weather and time of day) if outdoors, and the type of light source and window size if indoors. The influence of external light can be determined from the window size. The shooting range is, for example, information about the lens of the camera used (wide-angle or telephoto, focal length, etc.) and the installation position and angle of the camera. Image quality is the image quality of the entire image, such as the performance and set resolution of the camera used. Note that there is no particular limitation on the information about the background of the image. The usage condition information also relates to the attributes of people in the image, such as the characteristics of people who are expected to appear in the video (such as their age, sex, race, clothing, and posture).
[0027] The related information acquisition unit 802 searches the Web environment 105 or the user environment 101 for information about the background of the image and the attributes of the people in the image, and acquires the information of the Web page obtained as a result of the search as usage condition related information.
[0028] When obtaining the use condition related information from the Web environment 105 , the related information obtaining unit 802 obtains the use condition related information from the Web environment 105 using the search server 109 . The related information acquisition unit 802 performs a search using, for example, the function 602 "fall detection" as a query. As a result of the search, the related information acquisition unit 802 acquires, for example, article information about fall fraud from the news site 107. Assume that this article information includes text information such as "A fall fraud occurred. A white man in his 50s wearing a dress shirt fell over himself in the beverage section and sued the store." Furthermore, the related information acquisition unit 802 performs a search using, for example, "A super" of the order recipient 601 as a query. As a result of the search, the related information acquisition unit 802 acquires, for example, basic store information such as the layout of the sales floor and business hours from the user site 106 of A super. The related information acquisition unit 802 also performs a search using, for example, the "Las Vegas store" of the store used 604 as a query. As a result of the search, the related information acquisition unit 802 acquires, for example, the racial and age demographic composition of Las Vegas from the information site 108. Furthermore, the related information acquisition unit 802 performs a search using, for example, the model name "M30" of the camera 606 in use as a query. As a result of the search, the related information acquisition unit 802 acquires, for example, a review article about the M30 from the information site 108. Assume that this review article contains text information about image quality, such as "High image quality during the day, but high noise at night."
[0029] When obtaining usage condition related information from the user environment 101, the related information obtaining unit 802 obtains the usage condition related information from the user environment 101 using a search server (not shown) that can access the in-house mail server 102 and the sales data management server 103.
[0030] The related information acquisition unit 802 acquires testimony information from other stores from the in-house mail server 102, such as "A man in his 50s at the New York store complained that he fell in the vegetable section and was injured." Furthermore, the related information acquisition unit 802 acquires data on person attributes, such as "50% of the purchasing demographics are in their 40s, and 40% are in their 50s," from the sales data management server 103.
[0031] In S403, the image description generation unit 304 generates image description data based on the usage condition related information acquired in S402. The image description data is data for generating an image that is expected to be captured when using the video analysis device 104, and is described in text. In this embodiment, the image description data is generated according to a format defined for each video analysis function.
[0032] Below are examples of image description data templates for each video analysis function. 1.Fall detection function Generate an image with {quality} of a person of {race}, {age}, {gender}, wearing {clothing}, falling at {location} at {time}. 2. Gradient analysis function Generate an image with {quality} of a person of {race}, {age}, {gender}, wearing {clothing}, at {time} in {location}, stretching their hand. 3. Intrusion detection function Generate an image with {quality} of a person of {race}, {age}, {gender}, wearing {clothing}, entering the {area} of {location} at {time}. The race, age, etc. in the above {} represent image description items. The image description generation unit 304 extracts information corresponding to each image description item from the usage condition related information acquired in S402, and generates multiple variations of image description data using the extracted information and a template according to the video analysis function. Hereinafter, information corresponding to the image description items may be referred to as item information. Note that the image description items shown in 1 to 3 above for the image description data are examples, and are not particularly limited as long as they relate to the background of the image and the attributes of the subject (here, a person) to be analyzed.
[0033] Note that the image description data examples 1 to 3 above describe positive cases where the detection target is present, but depending on the function, image description data for negative cases where the detection target is not present may also be necessary. For example, in the case of the fall detection function (1), image description data for negative cases can be generated by changing the description "falling" to something else such as "walking" or "standing."
[0034] The process of generating image description data executed in S403 will be described in detail below with reference to the flowchart of FIG. In S501, the image description generation unit 304 acquires information corresponding to each image description item from the usage condition related information acquired in S402. Specifically, the image description generation unit 304 removes unnecessary information such as HTML tags from a web page or the like stored as usage condition related information in the secondary storage device 204, RAM 203, or the like, and extracts text information. Next, if the extracted text information is in Japanese, it is broken down into parts of speech using morphological analysis, and information corresponding to each image description item is extracted by matching it with a dictionary that further classifies nouns into categories such as race, age, and gender. Alternatively, information about each image description item may be extracted using a large-scale language model such as a Generative Pre-trained Transformer (GPT).
[0035] For example, from the aforementioned article information, "A fall fraud occurred. A white man in his 50s wearing a dress shirt fell over himself in the drinks section and sued the store," the following items of information about the person's attributes are acquired: age "50," clothing "dress shirt," race "white," and gender "male." In addition, the location of the photo, "the drinks section," is acquired as item information about the background of the image. Furthermore, from the review article "High image quality during the day, but high noise at night," the time "daytime" and image quality "high image quality," as well as the time "nighttime" and image quality "high noise" are acquired as item information related to the background of the image. Based on this, the image description generator 304 generates the following image description data: "Generating a high-quality image of a white male in his 50s wearing a dress shirt falling in a beverage aisle in daytime" "Generating a high-noise image of a white man in his 50s wearing a dress shirt falling in a drink aisle at night"
[0036] Next, in S502, the image description generation unit 304 refines the information corresponding to each image description item. Since the information corresponding to the image description items is extracted from the usage condition related information, it represents content that is close to the image that will actually be shot by the user. Therefore, it is desirable to ensure as many variations of images as possible that are generated from the information corresponding to the image description items. Therefore, the image description generation unit 304 increases the variation of image description data by refining the information (item information) corresponding to each image description item.
[0037] Specifically, the image description generation unit 304 adds information corresponding to sub-items to the item information. Sub-items are lower-level items for further detailing the item information. For example, if "drinks counter" is extracted as the "photography location," the image description generation unit 304 acquires the "area" within the counter as a sub-item and adds "in front of the shelf," "aisle," etc., corresponding to the "area." Furthermore, for example, if "shirt" is extracted as the "clothing," the image description generation unit 304 acquires "color" and "pattern" as sub-items of the shirt, and adds "white," "black," "red," etc., corresponding to the "color," and adds "solid," "striped," "checked," etc., corresponding to the "pattern." FIG. 7 shows an example of a table that defines sub-items. By referencing a table such as that shown in FIG. 7, the image description generation unit 304 acquires sub-items, which are lower-level items of the item information, and information corresponding to each sub-item.
[0038] For example, the image description generating unit 304 generates the following variations of image description data by refining the above image description data into the shooting location "drink counter" and the clothing "dress shirt." "Generating a high-resolution image of a Caucasian man in his 50s wearing a plain white shirt, slumped in front of a beverage shelf in the daytime." "Generating a high-resolution image of a Caucasian man in his 50s wearing a plain black shirt, slumped in front of a beverage shelf in the daytime." "Generating a high-quality image of a Caucasian man in his 50s wearing a plain red shirt, slumped in front of a beverage shelf in the daytime." "Generating a high-resolution image of a Caucasian man in his 50s wearing a white striped shirt, slumped in front of a beverage shelf in the daytime." (The rest is omitted.) As described above, by combining sub-items with information corresponding to each image description item and acquiring information corresponding to the sub-items, it is possible to increase the variety of image description data.
[0039] Here, the use condition related information includes information that is likely to have actually been photographed (highly credible), such as the testimonial information from other stores. The image description generation unit 304 may adjust the information obtained from such use condition related information with high importance so that more variations are generated. The importance here is determined, for example, depending on the source from which the use condition related information is obtained. The image description generation unit 304 may refine only the image description items obtained from testimonial information, or may increase the number of sub-items for image description items obtained from testimonial information compared to image description items obtained from sources other than testimonial information. In this way, the image description generation unit 304 adjusts the number of image description data generated based on the use condition related information according to the importance of the use condition related information.
[0040] Furthermore, for example, information on percentages such as "50% in their 40s, 40% in their 50s" may be extracted from the usage condition-related information acquired from the sales data management server 103 or the information site 108. In such a case, it is assumed that the people expected to appear in the video will have a similar percentage. Therefore, the image description generation unit 304 may adjust the percentages of image description data for ages "40s" and "50s" to 50% and 40%, respectively, of the total image description data to be generated.
[0041] Next, in S503, the image description generation unit 304 performs adjustments according to the image description items. Because the usage condition-related information represents typical content of the video that is actually shot, there is a bias, such as the inclusion of a large amount of specific information for certain image description items. However, some image description items may make it easier to detect people with specific attributes, such as race or gender, which can lead to ethical issues and other inconveniences. To address this, the image description generation unit 304 may perform adjustments so that the proportion of image description data in which a specific image description item is specific information does not exceed a threshold value. Specifically, if the proportion exceeds the threshold value, the generated image description data is randomly thinned out until it falls below the threshold value. The image description generating unit 304 stores the plurality of image description data generated as described above in the secondary storage device 204, RAM 203, etc. In this way, the series of processes for generating image description data is completed.
[0042] Returning to the explanation of Figure 4, in S404, the image generation unit 306 generates a training image (analysis image) based on the image description data generated in S403. Specifically, it uses Contrastive Language-Image Pre-training (CLIP), a model that is pre-trained so that text and images can be represented by common vectors. The image generation unit 306 encodes the image description data using CLIP as a text encoder, and uses a diffusion model or the like as conditioning when generating an image. This makes it possible to obtain an image that conforms to the image description data. Since there may be multiple images that are similar to the encoded image description data, multiple images can be obtained from a single piece of image description data.
[0043] Next, in S405, the image generation unit 306 generates correct labels for the learning images generated in S404 using image description data. This correct label is defined for each template of the image description data. For the image description data item 1 above, "falling," the correct label is defined as "falling," and for the negative examples, "walking" and "standing," the correct label is defined as "not falling." The generated image storage unit 307 assigns correct labels to the multiple generated images generated in S404 and stores them in the secondary storage device 204, RAM 203, or the like.
[0044] Next, in S406, the additional learning unit 309 reads out a trained video analysis model from the secondary storage device 204 or the like. The secondary storage device 204 stores video analysis models that have been trained in advance using existing images held by the developer for each video analysis function. The additional learning unit 309 reads out the video analysis model that corresponds to the function to be used (in this embodiment, fall detection).
[0045] In S407, the additional learning unit 309 performs additional learning of the video analysis model read out in S406 based on the generated image and correct label stored in S405. Specifically, the CPU 201 may use the weights of the trained model as initial values, fix the weights of the front part (backbone) that extracts basic features of the model, and update only the weights of the backbone. Of course, the method is not limited to the above, and any learning method that adapts the model to the generated image, such as updating the weights of the entire model, may be used.
[0046] In S408, the model providing unit 310 determines whether or not the user has started installing the video analysis device 104. Specifically, the model providing unit 310 checks this based on a notification from the user, such as an email. The process waits in S408 until the model providing unit 310 determines that installation of the video analysis device 104 has started. If it is determined that installation of the video analysis device 104 has started, the process proceeds to S409. In S409, the model providing unit 310 transmits the updated model that underwent additional learning in S407 to the video analysis device 104. For example, the video analysis device 104 may download the updated model from the learning device 111 via the Internet 112. The video analysis device 104 inputs the video captured by the camera into the video analysis model provided by the learning device 111 and performs video analysis. After that, the series of processes in the flowchart shown in FIG. 4 end.
[0047] According to the present embodiment, image description data representing video that is expected to be actually captured can be generated from information related to the user's usage conditions for the video analysis device, and learning images can be generated based on the image description data. This makes it possible to train a video analysis model using images similar to those seen when the user device is in operation, without requiring the user to provide video actually captured with the device. In other words, highly accurate video analysis can be performed from the start of operation of the user device.
[0048] <Embodiment 2> In the first embodiment, a method was described in which training images generated based on image description data are used directly for additional training. In the second embodiment, a method is described in which training images generated based on image description data are checked by a user, and feedback from the user is reflected before the images are used for additional training. This makes it possible to generate training images that are more suited to the usage environment. Note that a description of the parts common to the first embodiment will be omitted, and the differences from the first embodiment will be mainly described.
[0049] Figure 8 is a block diagram showing the functional configuration of a learning device 111 according to this embodiment. A usage condition input unit 801, a related information acquisition unit 802, a related information storage unit 803, and an image description generation unit 804 in Figure 8 correspond to the usage condition input unit 301, the related information acquisition unit 302, the related information storage unit 303, and the image description generation unit 304 in Figure 3. Furthermore, an image description storage unit 805, an image generation unit 806, and a generated image storage unit 807 in Figure 8 correspond to the image description storage unit 305, the image generation unit 306, and the generated image storage unit 307 in Figure 3. Furthermore, a trained model storage unit 808, an additional training unit 809, and a model providing unit 810 in Figure 8 correspond to the trained model storage unit 308, the additional training unit 309, and the model providing unit 310 in Figure 3.
[0050] A generated image providing unit 811 transmits the image generated by the image generating unit 806 to the user's PC 115 via the network I / F 207 . The feedback receiving unit 812 receives text information that is feedback input by the user on the PC 115 . The image description correcting unit 813 corrects the image description data stored in the image description storage unit 805 .
[0051] The basic processing flow executed by the learning device 111 according to this embodiment is the same as the flowchart shown in Fig. 4. Below, the processing flow of making corrections based on feedback from the user, which is a characteristic process of this embodiment, will be described with reference to Fig. 9.
[0052] 4, in S901, the generated image providing unit 811 selects a generated image to provide to the user. Since it is not realistic to have the user check all generated images, generated images are selected and provided. Note that the method of selecting images may be any method, such as randomly selecting a predetermined number of images, or selecting only images generated based on image description data in which predetermined image description items are specific information.
[0053] Next, in S902, generated image providing unit 811 transmits the generated image selected in S901 to PC 115 via Internet 112. Note that generated image providing unit 811 does not provide the generated image as is, but transmits HTML format data for displaying a confirmation UI screen having a screen configuration such as that shown in Fig. 10 on PC 115 in order to receive feedback from the user. In this way, generated image providing unit 811 presents to the user a confirmation UI screen displaying the generated image selected in S901.
[0054] FIG. 10 shows an example of the screen configuration of a confirmation UI screen presented to the user. The confirmation UI screen is displayed on a display device of PC 115. Images 1002 and 1003 (two images in this example) selected in S901 are displayed on confirmation UI screen 1001. Images 1002 and 1003 are images generated from image description data. To the right of images 1002 and 1003 are input fields 1004 and 1005 where the user can enter text comments. After the user enters feedback in input fields 1004 and 1005 via an input device of PC 115, when the user presses send button 1006, PC 115 transmits the contents (text information) of input fields 1004 and 1005 to learning device 111 via Internet 112.
[0055] Next, in S903, the feedback receiving unit 812 determines whether or not feedback (text information) has been received from the PC 115. The feedback receiving unit 812 waits in S903 until it receives feedback, and if it determines that the feedback receiving unit 812 has received feedback, the processing proceeds to S904.
[0056] In S904, the image description correction unit 813 corrects the image description data stored in the secondary storage device 204, RAM 203, etc., based on the text information received in S903. The image description correction unit 813 acquires information corresponding to each image description item from the received text information using the method described in S501. The image description correction unit 813 adds or replaces information to existing image description data using the information acquired from the text information.
[0057] In the example shown in FIG. 10, text information such as "The floor is a dark color" and "The lighting is brighter" is obtained as feedback. Based on this, information such as "Dark color floor" and "Bright lighting" can be added to the image description data to make it closer to the image that will actually be captured. Based on the image description data corrected in this manner, the processing from S404 onwards in FIG. 4 is executed. In other words, the image generation unit 306 regenerates the image using the image description data corrected according to the flowchart in FIG. 9.
[0058] According to the present embodiment, the image description data can be corrected by receiving user feedback on the generated images. This allows the video analysis model to be trained using images similar to those seen when the user device is in operation, without requiring the user to provide images actually taken with the user device. In other words, highly accurate video analysis can be performed from the very start of the user device's installation.
[0059] <Other embodiments> In the above-described embodiments, the image generation unit 306 generates an image based solely on image description data. However, an image may also be generated by using existing images posted on the news site 107, etc. For example, assume that the image description generation unit 304 acquires an image 1101, such as that shown in FIG. 11, from a web page stored as usage condition-related information. The image 1101 shows a person 1102 who has fallen. If the posture of the person who has fallen can be reflected in the training image, the accuracy of the fall detection function can be expected to improve. Therefore, the image description generation unit 304 acquires information on the joint points of the person 1102 (circles 1103 indicate joint points, and dotted lines indicate line segments connecting the joint points) using a method such as OpenPose. Next, the acquired information on the joint points is used as a condition to generate an image by using a method such as ControlNet so that the posture of the person in the image is similar. In this way, the image generation unit 306 generates an image that reflects the characteristics of the acquired image 1101. Here, since a fall detection function is assumed, the joint points of a person are acquired as image features, but the features acquired from the image are not particularly limited as long as they are features related to the video analysis function. For example, the image generation unit 306 may generate an image with a similar background using the results of dividing the image into units such as walls and floors. By generating an image using an existing image in this way, it is possible to generate an image with similar features to the image being used in combination.
[0060] A storage medium storing a program for realizing the functions of the learning device 111 according to each of the above-described embodiments may be supplied to a system or device, and a computer (or CPU, MPU, or GPU) of the system or device may read and execute the program code stored in the storage medium. Execution of the program code read by the computer not only realizes the functions of the above-described embodiments, but also includes an operating system (OS) running on the computer performing some or all of the actual processing based on the instructions of the program code.
[0061] Furthermore, the following method may be used to realize the functions of the learning device 111 according to the above-described embodiment. Program code read from a storage medium is written to memory on a function expansion card inserted into a computer or on a function expansion unit connected to the computer. Then, based on the instructions of the program code, a CPU or the like on the function expansion card or function expansion unit performs some or all of the actual processing. The storage medium stores program code corresponding to the flowcharts described above.
[0062] Although the present invention has been described above with reference to the embodiments, the above embodiments are merely illustrative of specific examples of how the present invention can be implemented, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.
[0063] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0064] The disclosure of each of the above-described embodiments includes the following configurations, methods, and programs. (Configuration 1) an acquisition means for acquiring related information related to the use conditions based on the use conditions of the video analysis; a description generating means for generating description data representing an image that is expected to be a target of the video analysis when the video analysis is performed based on the related information; an image generating means for generating an analysis image to be used in the video analysis based on the description data; An image generating device comprising: (Configuration 2) The image generating device described in configuration 1 is characterized in that the acquisition means acquires at least one of the following information as the usage conditions: the name of the store where the camera that captures the video to be analyzed is installed, the type of business in which the camera is used, and the location where the camera is installed, and acquires the related information based on the usage conditions. (Configuration 3) The image generating device described in configuration 1 is characterized in that the acquisition means acquires at least one of the type of camera that captures the video that is the target of the video analysis, the shooting resolution of the camera, the shooting time of the camera, and the type of video analysis as the usage conditions, and acquires the related information based on the usage conditions. (Configuration 4) The image generating device according to any one of configurations 1 to 3, characterized in that the description generating means acquires item information, which is information corresponding to a predetermined description item, from the related information, and generates the description data written in a format corresponding to the function of the video analysis based on the item information. (Configuration 5) 5. The image generating device according to configuration 4, wherein the predetermined description items are items relating to at least one of an object in the background of the image, lighting conditions, a photographing range, and image quality. (Configuration 6) 6. The image generating device according to configuration 4 or 5, wherein the predetermined description items are items that represent attributes of a subject to be analyzed. (Configuration 7) The image generating device described in any one of configurations 1 to 6, characterized in that the acquisition means searches a website for web pages related to the terms of use and acquires the related information from text information contained in the searched web pages. (Configuration 8) The image generating device according to configuration 7, wherein the image generating means generates the analysis image based on the description data and features extracted from existing images included in the searched web pages. (Configuration 9) 9. The image generating device according to any one of configurations 1 to 8, wherein the acquisition means acquires the related information from a server managed by a user who uses the video analysis. (Configuration 10) 10. The image generating device according to any one of configurations 1 to 9, further comprising a learning means for learning a model to be used in the video analysis using the analysis image generated by the image generating means. (Configuration 11) 7. The image generating device according to any one of configurations 4 to 6, wherein the description generating means increases the variation of the description data by detailing the item information. (Configuration 12) The image generating device described in any one of configurations 4 to 6 and 11, characterized in that the description generating means controls to adjust the ratio of the description data including specific item information to the entire generated description data. (Configuration 13) 13. The image generating device according to any one of configurations 1 to 12, wherein the description generating means controls to adjust the description data generated from the related information according to the importance of the related information. (Configuration 14) a presentation means for presenting the analysis image generated by the image generation means to a user who uses the video analysis; a receiving means for receiving feedback on the presented analysis image; a correction means for correcting the description data based on the feedback; 14. The image generating device according to any one of configurations 1 to 13, further comprising: (Configuration 15) 15. The image generating device according to configuration 14, wherein the presentation means presents a screen on which the analysis image and an input field for a comment on the analysis image are displayed. (Configuration 16) 11. The image generating device according to configuration 10, further comprising a transmitting means for transmitting the model learned by the learning means to a device that performs the video analysis in the usage environment. (Configuration 17) 17. The image generating device according to claim 10 or 16, wherein the model is a neural network. (Configuration 18) A video analysis system including a video analysis device that performs video analysis of a video captured by a camera, and a learning device that learns a model used in the video analysis, The learning device an acquisition means for acquiring related information relating to the use conditions of the video analysis device based on the use conditions of the video analysis device; a description generating means for generating description data representing an image that is expected to be a target of video analysis by the video analysis device based on the related information; an image generating means for generating an analysis image based on the description data; a learning means for learning a model used in the video analysis using the analysis image generated by the image generation means; and The video analysis device an analysis means for performing the video analysis using the model learned by the learning means; A video analysis system comprising: (method) an acquisition step of acquiring related information related to the use conditions based on the use conditions of the video analysis; a description generation step of generating description data representing an image that is expected to be a target of the video analysis when the video analysis is used based on the related information; an image generation step of generating an analysis image of a model to be used in the video analysis based on the description data; An image generating method comprising: (program) The computer of the image generating device, an acquisition means for acquiring related information related to the use conditions based on the use conditions of the video analysis; a description generating means for generating description data representing an image that is expected to be a target of the video analysis when the video analysis is performed based on the related information; an image generation means for generating an analysis image of a model to be used in the video analysis based on the description data; A program that functions as a [Explanation of symbols]
[0065] 102: In-house mail server, 103: Sales data management server, 104: Video analysis device, 106: User site, 107: News site, 108: Information site, 111: Learning device, 112: Internet
Claims
1. an acquisition means for acquiring related information related to the use conditions based on the use conditions of the video analysis; a description generating means for acquiring item information corresponding to predetermined description items from the related information, and generating description data based on the item information, which represents an image that is expected to be the subject of the video analysis when the video analysis is used, and which is described in accordance with the function of the video analysis; an image generating means for generating an analysis image to be used in the video analysis based on the description data; An image generating device comprising:
2. The image generating device described in claim 1, characterized in that the acquisition means acquires at least one of the following information as the usage conditions: the name of the store where the camera that captures the video to be analyzed is installed, the type of business in which the camera is used, and the location where the camera is installed, and acquires the related information based on the usage conditions.
3. The image generating device described in claim 1, characterized in that the acquisition means acquires at least one of the type of camera that captures the video to be analyzed, the camera's shooting resolution, the camera's shooting time, and the type of video analysis as the usage conditions, and acquires the related information based on the usage conditions.
4. 2. The image generating device according to claim 1, wherein the description generating means generates the description data written in a format corresponding to the function of the video analysis based on the item information.
5. 5. The image generating device according to claim 4, wherein the predetermined description items are items relating to at least one of an object in the background of the image, lighting conditions, a photographing range, and image quality.
6. 5. The image generating apparatus according to claim 4, wherein the predetermined description items are items that represent attributes of a subject to be analyzed.
7. 2. The image generating device according to claim 1, wherein the acquiring means searches a website for a web page related to the terms of use, and acquires the related information from text information included in the searched web page.
8. The image generating device according to claim 7, characterized in that the image generating means generates the analysis image based on the description data and features extracted from existing images contained in the searched web pages.
9. The image generating device according to claim 1 , wherein the acquisition means acquires the related information from a server managed by a user who uses the video analysis.
10. 2. The image generating device according to claim 1, further comprising: a learning unit that uses the analysis image generated by the image generating unit to learn a model used in the video analysis.
11. 5. The image generating apparatus according to claim 4, wherein the description generating means increases the variation of the description data by detailing the item information.
12. 5. The image generating apparatus according to claim 4, wherein said description generating means controls to adjust the ratio of said description data including said specific item information to the entirety of said generated description data.
13. 2. The image generating apparatus according to claim 1, wherein the description generating means adjusts the description data generated from the related information in accordance with the importance of the related information.
14. a presentation means for presenting the analysis image generated by the image generation means to a user who uses the video analysis; a receiving means for receiving feedback on the presented analysis image; a correction means for correcting the description data based on the feedback; 2. The image generating device of claim 1, further comprising:
15. 15. The image generating device according to claim 14, wherein the presenting means presents a screen on which the analysis image and an input field for a comment on the analysis image are displayed.
16. 11. The image generating device according to claim 10, further comprising a transmitting unit that transmits the model learned by the learning unit to a device that performs the video analysis in the environment in which the device is used.
17. 11. The image generating apparatus of claim 10, wherein the model comprises a neural network.
18. A video analysis system including a video analysis device that performs video analysis of a video captured by a camera, and a learning device that learns a model used in the video analysis, The learning device an acquisition means for acquiring related information relating to the use conditions of the video analysis device based on the use conditions of the video analysis device; a description generating means for acquiring item information corresponding to predetermined description items from the related information, and generating description data based on the item information, which represents images that are expected to be the subject of video analysis by the video analysis device, and which is described in accordance with the function of the video analysis; an image generating means for generating an analysis image based on the description data; a learning means for learning the model using the analysis image generated by the image generation means; and The video analysis device an analysis means for performing the video analysis using the model learned by the learning means; A video analysis system comprising:
19. an acquisition step of acquiring related information related to the use conditions based on the use conditions of the video analysis; a description generation step of acquiring item information corresponding to predetermined description items from the related information, and generating description data based on the item information, which represents an image that is expected to be the subject of the video analysis when the video analysis is used, and which is described in accordance with the function of the video analysis; an image generation step of generating an analysis image to be used in the video analysis based on the description data; An image generating method comprising:
20. The computer of the image generating device, an acquisition means for acquiring related information related to the use conditions based on the use conditions of the video analysis; a description generating means for acquiring item information corresponding to predetermined description items from the related information, and generating description data based on the item information, which represents an image that is expected to be the subject of the video analysis when the video analysis is used, and which is described in accordance with the function of the video analysis; an image generation means for generating an analysis image to be used in the video analysis based on the description data; A program that functions as a
Citation Information
Patent Citations
Learned model provision method and learned model provision device
WO2018142766A1
Monitoring system, analyzing device, and ai model generating method
WO2022059122A1