Information processor and information processing method
The information processing apparatus enhances the accuracy of state estimation by using a determination unit to ensure that only complete detection targets are analyzed using VQA, addressing the issues of over-detection and partial image cuts in conventional techniques.
Patent Information
- Application Number
- JP2023206158
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-18
AI Technical Summary
Conventional techniques for estimating specific states using VQA struggle with accuracy, particularly when the detection target is partially cut off in the captured image, leading to over-detection and insufficient estimation.
An information processing apparatus that includes a determination unit to assess whether the imaging situation of a detection target meets a predetermined condition, and an execution unit that executes VQA processing only when the condition is satisfied, ensuring that only complete detection targets within a defined range are subjected to VQA analysis.
This approach enables highly accurate state estimation by excluding partially cut-off or over-detected targets from the VQA judgment, thereby improving the reliability of the estimation process.
Smart Images

Figure 2025091110000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus and an information processing method.
Background Art
[0002] Conventionally, a technique for estimating a specific state based on predetermined information has been known. For example, there is a known technique for estimating an equipment state or a dangerous state that violates a safety manual based on a captured image (on-site image) by using an AI technique called VQA (Visual Question Answering).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the conventional technique, there is a case where the state is estimated even when a sufficient state cannot be estimated, so there is room for further improvement for performing highly accurate estimation.
[0005] The present application has been made in view of the above, and an object thereof is to provide an information processing apparatus and an information processing method that enable highly accurate estimation.
Means for Solving the Problems
[0006] The information processing apparatus according to the present application includes a determination unit that determines whether an imaging situation of a detection target satisfies a predetermined condition, and an execution unit that executes processing using VQA (Visual Question Answering) when the determination unit determines that the predetermined condition is satisfied.
Effects of the Invention
[0007] According to one aspect of the embodiment, it is possible to achieve the effect of enabling highly accurate estimation.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
[0009] Hereinafter, embodiments for carrying out the information processing apparatus and the information processing method according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing apparatus and the information processing method according to the present application are not limited by this embodiment. Also, in the following embodiments, the same parts are denoted by the same reference numerals, and duplicate explanations are omitted.
[0010] Conventionally, a technique for estimating a specific state based on predetermined information is known. For example, there is a known technique for estimating an equipment state or a dangerous state that violates a safety manual based on a captured image by using an AI technique called VQA. VQA is a technique that integrates computer vision and natural language processing, accepts questions about objects and scenes in a captured image in natural language, and generates answers thereto.
[0011] VQA is a model that generates an answer to a question based on an image when an image and a question are input. For example, when an image of a worker wearing gloves and a question "Are you wearing gloves?" are input, VQA outputs "YES" based on the image. Using such VQA technology, needs such as detecting violations of safety manuals and the occurrence of highly urgent situations can be considered. For example, there is a need to detect a dangerous state that violates a safety manual based on a captured image captured by an imaging device (for example, a camera) installed at a specific site or the like.
[0012] However, in the prior art, there is room for further improvement to perform highly accurate estimation because the state may be estimated even when it cannot be sufficiently estimated. For example, in the prior art, when the detection target is only partially reflected at the edge of the captured image to be judged by VQA (for example, when the detection target at the edge of the captured image is cut off due to the angle of view of imaging), etc., when the state of the detection target cannot be sufficiently estimated (that is, when the state of the detection target in the captured image is not appropriate as the judgment target of VQA), the state may be estimated.
[0013] More specifically, when the detection target is an operator and the content to be inquired of VQA is something like "Is the operator wearing gloves?", there may be a case where it is impossible to appropriately determine whether the operator is wearing gloves, such as when the operator's hand is cut off at the edge of the captured image. Also, for example, since the technology for detecting an object or a person from a captured image uses a detection technology for detecting what is where in the captured image, it can be considered that there may be a case where highly accurate estimation cannot be performed due to over-detection in which more objects or people are detected than actually exist.
[0014] In the following embodiment, by specifying in advance the analysis range of the captured image that is the target of state estimation of the detection target, when it is impossible to sufficiently estimate the state, such as when the detection target is cut off, state estimation is not performed, enabling highly accurate estimation.
[0015] (Embodiment) 〔Configuration of Information Processing System〕 The information processing system 1 shown in FIG. 1 will be described. As shown in FIG. 1, the information processing system 1 includes a terminal device 10 and an information processing device 100. The terminal device 10 and the information processing device 100 are communicably connected by wire or wirelessly via a predetermined communication network (network N). FIG. 1 is a diagram showing a configuration example of the information processing system 1 according to the embodiment.
[0016] The terminal device 10 is an information processing device used by a user who desires to confirm the state of a detection target (it may also be to confirm the occurrence of a predetermined situation). The user is, for example, a manager who manages whether a worker working at a predetermined site (such as an employee of the user) is in a dangerous state (such as whether the equipment state violates the safety manual) (the user is not particularly limited to such examples). The terminal device 10 may be any device as long as it can realize the processing in the embodiment. The terminal device 10 may be, for example, a device such as a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, or a PDA. FIG. 10 shows the case where the terminal device 10 is a smartphone.
[0017] The terminal device 10 is, for example, a smart device such as a smartphone or smart glasses, and is a portable terminal device that can communicate with an arbitrary server device via a wireless communication network such as 4G to 5G (Generation) or LTE (Long Term Evolution). Further, the terminal device 10 has a screen such as a liquid crystal display, and has a screen having a touch panel function, and may receive various operations on display data such as content, such as a tap operation, a slide operation, and a scroll operation, by a finger or a stylus from the user. In FIG. 10, the terminal device 10 is used by the user U1.
[0018] The information processing apparatus 100 is an information processing apparatus aimed at enabling highly accurate estimation, and it may be any apparatus as long as it can realize the processing in the embodiment. For example, the information processing apparatus 100 may be an information processing apparatus aimed at enabling highly accurate estimation using VQA in order to meet the user's desire to check the state of the detection target. For example, the information processing apparatus 100 enables highly accurate estimation using VQA by determining whether the captured image is the image to be judged by VQA. Specifically, the information processing apparatus 100 enables highly accurate estimation using VQA by setting only the captured images that meet a predetermined condition as the objects to be judged by VQA. The information processing apparatus 100 is realized, for example, by a server apparatus or a cloud system that manages a predetermined site (for example, manages the safety of the site, etc.). For example, it is realized by a server apparatus or a cloud system such as a security system. The information processing apparatus 100 determines, for example, whether the imaging situation of the detection target meets a predetermined condition, and executes processing using VQA when it is determined that the predetermined condition is met. As a result, for example, highly accurate estimation regarding the state of the detection target becomes possible.
[0019] Note that in FIG. 1, the case where the terminal device 10 and the information processing apparatus 100 are separate devices is shown, but the terminal device 10 and the information processing apparatus 100 may be integrated.
[0020] In VQA, since multiple questions are asked simultaneously, if the detection target is only partially reflected at the edge of the captured image to be judged by VQA, even if it is possible to determine whether the operator is standing in response to the question "Are you standing?", if the operator's hand is not reflected, it may not be possible to determine whether the operator is wearing gloves in response to the question "Are you wearing gloves?". Thus, when multiple questions are asked simultaneously, there may be a question that causes problems in the VQA judgment. Therefore, in the following embodiments, the detection target in which the whole body (entirety) is reflected is used as the VQA judgment target. FIG. 2 is an explanatory diagram for explaining the judgment target according to the embodiment. Specifically, FIG. 2 is a diagram for judging whether the detection target B, which is a person, is wearing gloves, and is a diagram showing a captured image captured from the ceiling of the work space of the detection target B. In the upper part of FIG. 2, the detection target B is cut off from the edge of the captured image, and a part (such as a hand) of the detection target B is not reflected. Therefore, since VQA cannot appropriately answer the question, a captured image in which a part of such a detection target is cut off is excluded from the VQA judgment target. On the other hand, in the lower part of FIG. 2, the detection target B is contained within the captured image and the whole body is reflected. Therefore, VQA can appropriately answer the question. A captured image in which the whole of such a detection target is reflected corresponds to the VQA judgment target.
[0021] In the following embodiments, an imaging image is acquired from an imaging device located at a predetermined height. In the prior art, since the viewing angle changes depending on the installation location, the size of an object or a person can vary significantly. As a premise of the following embodiments, for example, by attaching an imaging device to the ceiling such as ViewLED (LED lighting with a camera), it becomes possible to realize a consistent viewing point. FIG. 3 is an explanatory diagram for explaining the imaging viewing point according to the embodiment. On the left side of FIG. 3, the viewing point is at a height of 8 m, and since it is imaged from a height of 8 m, the apparent size is smaller than that in the case of the right side of FIG. 3. On the other hand, on the right side of FIG. 3, the viewing point is at a height of 2 m, and since it is imaged from a height of 2 m, the apparent size is larger than that in the case of the left side of FIG. 3. On the other hand, in the middle of FIG. 3, the viewing point is at a height of 4 m, and since it is imaged from a height of 4 m, the apparent size is in the middle between the cases of the left side and the right side of FIG. 3. Thus, the apparent size changes according to the distance. Also, if installations other than the ceiling are not assumed, the variable that determines the apparent size is only the height.
[0022] As a premise of the following embodiments, for example, the apparent size is determined based on, as a variable, the distance between the viewpoint and the detection target, angle information, lens information (e.g., focal length of the lens, etc.), sensor information (e.g., sensor size, etc.), and the like. For example, it is determined based on a coefficient corresponding to the angle information, lens information, and sensor information, and a variable of the distance between the viewpoint and the detection target. At this time, for example, it is determined based on a calculation formula having the variable of the distance between the viewpoint and the detection target as the denominator. Note that the calculation formulas for determining the apparent size may be different between a normal lens (central projection method) and a fish-eye lens (equidistant projection method). Further, in the following embodiments, the positional relationship between the viewpoint and the detection target is not limited to the case where the detection target is directly below the imaging device, as in the example of FIG. 3. FIG. 4 is an explanatory diagram for explaining the imaging viewpoint according to the embodiment. As in the example of FIG. 4, even when the detection target B is not directly below the imaging device, since the position of the imaging device does not change only because the distance to the detection target B changes (the change is only in the height direction), it is possible to determine the apparent size for each height. On the other hand, in the prior art, since the installation position of the imaging device is not determined and the variable is not only the height, it is difficult to determine the apparent size.
[0023] In the following embodiments, a threshold value for the detection target is set in advance so that the detection target is not missed at the edge of the captured image. For example, the threshold value is set based on the height and a coefficient corresponding to the imaging device. When the detection target overlaps a predetermined range (distance from the edge of the captured image) at the edge of the captured image corresponding to the set threshold value, it is excluded from the determination target of VQA. FIG. 5 is an explanatory diagram for explaining the determination target range according to the embodiment. As in the example of FIG. 5, if even a part of the rectangle indicating the detection target is outside the determination target range, it is regarded that the detection target is missed and excluded from the determination target. The left side of FIG. 5 is an explanatory diagram for explaining the determination target range when the threshold value is set at the edge of the captured image, and the right side of FIG. 5 is an explanatory diagram for explaining the detection target that becomes the determination target and the detection target that does not become the determination target. FIG. 6 is an explanatory diagram for explaining the determination target range according to the embodiment using an actual example. In the upper part of FIG. 6, since a part of the rectangle indicating the detection target overlaps the predetermined range at the edge of the captured image, it is excluded from the determination target. On the other hand, in the lower part of FIG. 6, since none of the rectangles indicating the detection target overlaps the predetermined range at the edge of the captured image, it becomes the determination target. On the other hand, as in the prior art, when the installation position of the imaging device is not determined, since the size of the rectangle also varies according to the installation position, the threshold value cannot be set.
[0024] In the following embodiments, only the captured images in which the entire body of an object or person to be detected is shown are extracted so that over-detection in which more objects or people are detected than in reality does not occur. As a technique for detecting an object or person to be detected from a captured image, for example, a detection technique for detecting what is where in the captured image is used. Here, an example of a method for reducing over-detection will be described. A threshold value is set for the coordinate data of the rectangle of the object to be detected in the captured image. The threshold value is set for the purpose of excluding an object or person whose body part at the edge of the captured image is cut off. For example, it is set according to the distance from the edge of the captured image. When the threshold value is set within the predetermined range of the dotted line on the left side of FIG. 5, the three rectangles at the edge of the captured image are excluded from the determination target, and one rectangle near the center becomes the determination target. If even one of the coordinate data of the rectangle (for example, the coordinate data of the four corners) does not satisfy the threshold value, it is excluded from the determination target. In addition, as an example of a method for setting the threshold value, in other cases, even if an object is located at the center of the captured image but the whole body is not shown because it is blocked by a body, or if it is too small for the AI to make a determination, it can be excluded from the determination target. For example, it is also possible to set the threshold value for the length of one side of the rectangle obtained from the coordinate data to exclude it according to the size of the rectangle. In this way, by setting the threshold value for the coordinate data or the size of the rectangle, it is possible to reduce over-detection. In addition, for example, it is also possible to detect the skeleton after detecting a person and exclude it from the determination target if not all the skeleton points are included. Also, it may be determined whether the range corresponding to the question text of VQA satisfies a predetermined condition, and if the predetermined condition is not satisfied, it may be excluded from the determination target. For example, when the description of "person" is included in the question text of VQA, "person" is detected, and it is determined whether the detected range of "person" satisfies a predetermined condition. If the predetermined condition is satisfied, it is set as the determination target, and if the predetermined condition is not satisfied, it may be excluded from the determination target.
[0025] In the following embodiments, by acquiring the coordinate data of a rectangle indicating a detection target, if even one of the coordinates of the vertices of the rectangle is outside the determination target range, it is determined that the target is missing and excluded from the determination target. Further, the determination target range is determined based on the thresholds for the vertical (y-axis of the coordinates) and horizontal (x-axis of the coordinates) of the captured image. Further, since the threshold is determined by an equation such that the value decreases as the ceiling height increases, the determination target range becomes wider as the ceiling height increases. FIG. 7 is an explanatory diagram for explaining the threshold according to the embodiment. In FIG. 7, axis thresholds are set for each of the vertical and horizontal directions of the captured image. For example, in FIG. 7, with the upper left of the captured image as the origin of the axis (x = 0 and y = 0), thresholds are set at positions x = 55 and x = 1225 of the captured image (note that the number of pixels on the x-axis is 1280 and the right end of the captured image is x = 1280). The region where x > 55 and x < 1225 of the captured image is the determination target range, and the regions where x < 55 and x > 1225 of the captured image are outside the determination target range. Also, thresholds are set at positions y = 5 and y = 775 of the captured image (note that the number of pixels on the y-axis is 780 and the lower end of the captured image is y = 780). The region where y > 5 and y < 775 of the captured image is the determination target range, and the regions where y < 55 and y > 775 of the captured image are outside the determination target range.
[0026] In the following embodiments, captured images of detection targets with a rectangle size equal to or less than a predetermined threshold are excluded from the determination target. Since the size of the rectangle is determined in the captured image from the ceiling viewpoint, captured images of detection targets that are too small or that only partially appear are excluded from the determination target. For example, captured images of detection targets where the region on the x-axis does not satisfy 100 pixels (detection targets where x2 - x1 < 100) or the region on the y-axis does not satisfy 100 pixels (y2 - y1 < 100) are excluded from the determination target. FIG. 8 is an explanatory diagram for explaining the relationship with the detection target when a threshold is set for the size of the rectangle according to the embodiment. Further, FIG. 9 is an explanatory diagram for explaining whether a captured image of a detection target becomes a determination target when the size of the rectangle according to the embodiment is applied to the detection target. In FIG. 9, captured images of detection targets where at least one of the regions on the x-axis and the region on the y-axis is less than 100 pixels are excluded from the determination target.
[0027] In the following embodiments, by attaching an imaging device to the ceiling, a consistent viewpoint can be realized, enabling the setting of thresholds. In the prior art, the angle of view changed depending on the installation location, and the size of objects and people fluctuated significantly, making it difficult to set thresholds based on coordinates. However, in the following embodiments, since the viewing angle (the size of objects and people, etc.) is determined by the distance between the viewpoint and the detection target, the threshold can be uniquely determined. Also, in the following embodiments, since settings other than the ceiling are not assumed, the only variable for determining the size is the height. For this reason, in the prior art, it was necessary for the user to set each device individually for threshold input, but in the following embodiments, since the imaging device is attached to the ceiling, the threshold can be set without user intervention.
[0028] 〔Example of Information Processing〕 FIG. 10 is a diagram showing an example of information processing of the information processing system 1 according to the embodiment. The imaging device C is an imaging device installed on the ceiling. Hereinafter, FIG. 10 will be described with appropriate reference to FIGS. 2 to 9.
[0029] The information processing apparatus 100 acquires a captured image captured by the imaging device C (since the imaging device C is an imaging device installed on the ceiling, it is a captured image captured from a predetermined height position) (step S101). For example, the information processing apparatus 100 acquires the captured image on the left side of FIG. 2 or the captured image on the left side of FIG. 6. Then, the information processing apparatus 100 acquires rectangular coordinate data indicating the detection target from the acquired captured image (step S102). For example, the information processing apparatus 100 acquires the coordinate data on the left side of FIG. 5 or the coordinate data of FIG. 7. At this time, the captured image is not limited to a still image and may be a moving image.
[0030] The information processing apparatus 100 determines whether the imaging situation of the detection target satisfies a predetermined condition by determining whether the acquired coordinate data satisfies a predetermined condition (step S103). For example, the information processing apparatus 100 determines whether the coordinate data is within the determination target range (whether the detection target does not overlap with a predetermined range outside the determination target range) (see FIGS. 5 and 6). When the information processing apparatus 100 determines that the coordinate data is within the determination target range (when it is determined that the detection target does not overlap with a predetermined range outside the determination target range), it executes processing using VQA (step S104). That is, the information processing apparatus 100 sets the captured image of the detection target for which the coordinate data is determined to be within the determination target range as the determination target. On the other hand, when the information processing apparatus 100 determines that the coordinate data is not within the determination target range (when it is determined that the detection target overlaps with a predetermined range outside the determination target range), it excludes the captured image from the determination target and stores the excluded captured image. At this time, the information processing apparatus 100 may provide the terminal device 10 with information indicating that the captured image has been excluded from the determination target.
[0031] The information processing apparatus 100 acquires in advance a question and an assumed answer for the question, and in step S104, determines the similarity between the assumed answer and the estimated answer by acquiring (or generating) the estimated answer from the captured image using VQA (step S105). Then, when the similarity exceeds a predetermined threshold, the information processing apparatus 100 determines that there is no abnormality, and when the similarity does not exceed the predetermined threshold, it determines that there is an abnormality. At this time, the information processing apparatus 100 may provide the terminal device 10 with the result information (step S106). For example, when it is determined that there is no abnormality, the information processing apparatus 100 may provide information regarding the estimated answer. For example, when the estimated answer to the question "Are you wearing gloves?" is "YES" or "wearing gloves", the information processing apparatus 100 may provide information indicating that the detection target is "wearing gloves". Also, for example, when it is determined that there is an abnormality, the information processing apparatus 100 may provide information indicating that there is an abnormality in the estimated answer (for example, that the accuracy of the estimated answer is low or that information regarding the estimated answer cannot be provided).
[0032] 〔Configuration of the Terminal Device〕 Next, with reference to FIG. 11, the configuration of the terminal device 10 according to the embodiment will be described. FIG. 11 is a diagram showing a configuration example of the terminal device 10 according to the embodiment. As shown in FIG. 11, the terminal device 10 includes a communication unit 11, an input unit 12, an output unit 13, and a control unit 14.
[0033] (Communication Unit 11) The communication unit 11 is realized by, for example, a NIC (Network Interface Card) or the like. Then, the communication unit 11 is connected to a predetermined network N by wire or wirelessly, and information is transmitted and received between the communication unit 11 and an information processing device 100 or the like via the predetermined network N.
[0034] (Input Unit 12) The input unit 12 receives various operations from the user. In FIG. 10, the input unit 12 receives various operations from the user U1. For example, the input unit 12 may receive various operations from the user via the display surface by the touch panel function. Further, the input unit 12 may receive various operations from buttons provided on the terminal device 10 or from a keyboard or a mouse connected to the terminal device 10.
[0035] (Output Unit 13) The output unit 13 is a display screen of a tablet terminal or the like realized by, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and is a display device for displaying various information. The output unit 13 displays, for example, information received from the information processing device 100. For example, the output unit 13 displays information indicating the state of the detection target in response to a question input by the user in VQA. For example, the output unit 13 displays information regarding the estimated answer generated by VQA.
[0036] (Control Unit 14) The control unit 14 is, for example, a controller, which is realized by various programs stored in the storage device inside the terminal device 10 being executed with the RAM (Random Access Memory) as the working area by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like. For example, these various programs include the programs of applications installed in the terminal device 10. For example, these various programs include the programs of applications for displaying the information received from the information processing device 100. Further, the control unit 14 is realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0037] As shown in FIG. 11, the control unit 14 has a receiving unit 141 and a transmitting unit 142, and realizes or executes the operations of information processing described below.
[0038] (Receiving Unit 141) The receiving unit 141 receives, for example, information regarding the state of the detection target. For example, the receiving unit 141 receives information regarding whether it is possible to estimate the state of the detection target from the captured image (information regarding whether there is an abnormality) with respect to the question input by the user in VQA. Further, the receiving unit 141 receives, for example, information regarding whether the captured image has been excluded from the determination target (information regarding whether it can be a target of VQA).
[0039] (Transmitting Unit 142) The transmitting unit 142 transmits, for example, information regarding the captured image designated by the user when the user designates a captured image that the user wants to make the target of VQA. Further, the transmitting unit 142 transmits, for example, information regarding the question input by the user in VQA.
[0040] 〔Configuration of Information Processing Device〕 Next, with reference to FIG. 12, the configuration of the information processing apparatus 100 according to the embodiment will be described. FIG. 12 is a diagram showing a configuration example of the information processing apparatus 100 according to the embodiment. As shown in FIG. 12, the information processing apparatus 100 includes a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing apparatus 100 may include an input unit (for example, a keyboard or a mouse) that receives various operations from the administrator of the information processing apparatus 100, and a display unit (for example, a liquid crystal display) that displays various information.
[0041] (Communication unit 110) The communication unit 110 is realized by, for example, a NIC or the like. Then, the communication unit 110 is connected to the network N by wire or wirelessly, and information is transmitted and received to and from the terminal device 10 or the like via the network N.
[0042] (Storage unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in FIG. 12, the storage unit 120 includes an assumed answer storage unit 121 and a VQA model storage unit 122.
[0043] The assumed answer storage unit 121 stores information regarding assumed answers preset as answers to questions. The assumed answer storage unit 121 stores, for example, questions and their corresponding assumed answers in association with each other. Here, FIG. 13 shows an example of the assumed answer storage unit 121 according to the embodiment. As shown in FIG. 13, the assumed answer storage unit 121 has items such as "question answer ID", "question", and "assumed answer".
[0044] "Question-Answer ID" indicates identification information for identifying a combination of a question and an assumed answer. "Question" indicates a question (the content of the question). In the example shown in FIG. 13, an example where conceptual information such as "Question #1" or "Question #2" is stored in "Question" is shown, but actually, the text information of the question is stored. "Assumed Answer" indicates an assumed answer (the content of the assumed answer). In the example shown in FIG. 13, an example where conceptual information such as "Assumed Answer #1" or "Assumed Answer #2" is stored in "Assumed Answer" is shown, but actually, the text information of the assumed answer is stored.
[0045] The VQA model storage unit 122 stores information related to the VQA model. Here, FIG. 14 shows an example of the VQA model storage unit 122 according to the embodiment. As shown in FIG. 14, the VQA model storage unit 122 has items such as "VQA model ID" and "VQA model".
[0046] "VQA model ID" indicates identification information for identifying the VQA model. "VQA model" indicates the VQA model. In the example shown in FIG. 14, an example where conceptual information such as "VQA model #1" or "VQA model #2" is stored in "VQA model" is shown, but actually, the teacher data of the VQA model is stored.
[0047] (Control unit 130) The control unit 130 is a controller and is realized, for example, by various programs stored in the storage device inside the information processing device 100 being executed with the RAM as a work area by a CPU, MPU, etc. Also, the control unit 130 is realized by an integrated circuit such as an ASIC or FPGA.
[0048] As shown in FIG. 12, the control unit 130 has an acquisition unit 131, a determination unit 132, an execution unit 133, and a provision unit 134, and realizes or executes the information processing operations described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 12, and other configurations may be used as long as they can perform the information processing described later.
[0049] (Acquisition unit 131) The acquisition unit 131 acquires various types of information from the storage unit 120. Further, the acquisition unit 131 stores the acquired various types of information in the storage unit 120.
[0050] The acquisition unit 131 acquires various types of information from an external information processing device. The acquisition unit 131 acquires various types of information from other information processing devices such as the terminal device 10.
[0051] The acquisition unit 131 acquires, for example, a captured image (which may be a still image or a moving image). For example, the acquisition unit 131 acquires information regarding the captured image specified by the user. For example, the acquisition unit 131 acquires information regarding a captured image captured from a position at a predetermined height (for example, the position of the ceiling).
[0052] The acquisition unit 131 acquires, for example, coordinate data of a rectangle indicating a detection target. For example, the acquisition unit 131 acquires coordinate data of the vertices of the rectangle (for example, the four corners of the rectangle). For example, the acquisition unit 131 acquires coordinate data with a part of the area of the captured image as the origin. Further, the acquisition unit 131 acquires, for example, information regarding the length of the sides of the rectangle based on the acquired coordinate data.
[0053] The acquisition unit 131 acquires, for example, information regarding a threshold value preset according to the imaging position of the captured image. Further, the acquisition unit 131 acquires, for example, information (for example, coordinate data of a predetermined range) regarding a predetermined range (for example, the distance from the edge of the captured image) according to the acquired threshold value. Further, the acquisition unit 131 acquires, for example, information (for example, coordinate data of the determination target range) regarding the determination target range.
[0054] The acquisition unit 131 acquires, for example, information regarding a question and an assumed answer to the question. Further, the acquisition unit 131 acquires, for example, information regarding the user's question, and acquires information regarding an estimated answer generated using VQA from the captured image and the user's question.
[0055] (Determination unit 132) The determination unit 132 determines, for example, whether the coordinate data of the detection target acquired by the acquisition unit 131 satisfies a predetermined condition. In other words, the determination unit 132 determines, for example, whether the imaging situation of the detection target satisfies a predetermined condition. For example, the determination unit 132 determines whether all of the coordinate data of the detection target is within the determination target range (that is, whether the detection target does not overlap even partially with a predetermined range outside the determination target range).
[0056] (Execution unit 133) For example, when the determination unit 132 determines that all of the coordinate data of the detection target is within the determination target range, the execution unit 133 executes processing using VQA. For example, the execution unit 133 sets the captured image of the detection target for which the coordinate data has been determined to be within the determination target range as the determination target. Also, for example, when the determination unit 132 determines that the coordinate data of the detection target is not within the determination target range, the execution unit 133 excludes the captured image of the detection target for which the coordinate data has been determined not to be within the determination target range from the determination target. At this time, the execution unit 133 may store, for example, information regarding the captured image excluded from the determination target.
[0057] As an example of processing using VQA, the execution unit 133 determines, for example, the similarity between the estimated answer generated using VQA from the captured image acquired by the acquisition unit 131 and the user's question, and the assumed answer predetermined for that question. For example, when the similarity between the estimated answer and the assumed answer exceeds a predetermined threshold, the execution unit 133 determines that there is no abnormality, and when the similarity between the estimated answer and the assumed answer does not exceed the predetermined threshold, the execution unit 133 determines that there is an abnormality.
[0058] (Provision unit 134) For example, when the determination unit 132 determines that the coordinate data of the detection target is not within the determination target range, the provision unit 134 provides information indicating that the execution unit 133 has excluded it from the determination target. In other words, the provision unit 134 provides, for example, information indicating that the captured image has been excluded from the determination target because the coordinate data of the detection target is not within the determination target range. For example, the provision unit 134 provides information indicating that the detection target is missing and thus not subject to VQA determination.
[0059] When the similarity is determined by the execution unit 133 to exceed a predetermined threshold (that is, when it is determined that there is no abnormality), for example, the providing unit 134 provides information regarding the estimated answer generated using VQA. Further, for example, when the similarity is determined by the execution unit 133 not to exceed a predetermined threshold (that is, when it is determined that there is an abnormality), the providing unit 134 provides information indicating that an answer cannot be provided, information indicating that the accuracy of the answer is low, and the like. For example, when the estimated answer to the question "Are you wearing gloves?" is "NO" and the assumed answer is "YES", the providing unit 134 provides information such as "Either possibility is there". Further, for example, when the estimated answer to the question "Are you wearing gloves?" is "PROBABLY" and the assumed answer is "YES", the providing unit 134 provides information such as "There is a high possibility of wearing gloves".
[0060] 〔Flow of information processing〕 Next, with reference to FIGS. 15 and 16, the procedure of information processing by the information processing system 1 according to the embodiment will be described. FIGS. 15 and 16 are flowcharts showing the procedure of information processing by the information processing system 1 according to the embodiment.
[0061] As shown in FIG. 15, the information processing apparatus 100 acquires a captured image of the detection target (step S201).
[0062] The information processing apparatus 100 acquires coordinate data of the detection target from the captured image of the detection target (step S202).
[0063] The information processing apparatus 100 determines whether the coordinate data of the detection target satisfies a predetermined condition (step S203).
[0064] When it is determined that the coordinate data of the detection target satisfies a predetermined condition (step S203; YES), the information processing apparatus 100 executes processing using VQA (step S204). In this case, the information processing apparatus 100 executes the information processing shown in FIG. 16.
[0065] When it is determined that the coordinate data of the detection target does not satisfy a predetermined condition (step S203; NO), the information processing apparatus 100 excludes the captured image from the determination target (step S205). Then, the information processing apparatus 100 stores the captured image excluded from the determination target (step S206).
[0066] As shown in FIG. 16, the information processing apparatus 100 acquires a question and an assumed answer (step S301).
[0067] The information processing apparatus 100 acquires a captured image in which the coordinate data of the detection target satisfies a predetermined condition and a user's question (step S302). Specifically, the information processing apparatus 100 acquires the captured image determined to satisfy the predetermined condition in FIG. 15.
[0068] The information processing apparatus 100 generates an estimated answer from the captured image and the user's question using VQA (step S303).
[0069] The information processing apparatus 100 determines whether the similarity between the estimated answer and the assumed answer exceeds a predetermined threshold (step S304).
[0070] When it is determined that the similarity between the estimated answer and the assumed answer exceeds a predetermined threshold (step S304; YES), the information processing apparatus 100 determines that there is no abnormality (step S305). On the other hand, when it is determined that the similarity between the estimated answer and the assumed answer does not exceed a predetermined threshold (step S304; NO), the information processing apparatus 100 determines that there is an abnormality (step S306).
[0071] 〔Other system configuration examples〕 Note that the configuration of the information processing system 1 described above is merely an example, and the information processing system 1 can adopt any device configuration as long as the desired processing is possible. For example, the information processing system 1 may have a configuration without the terminal device 10. Also, various processes performed by the information processing device 100 may be executed not only by the information processing device 100 but also by any device included in the information processing system 1. For example, at least a part of various processes performed by the information processing device 100 (such as the determination unit 132, the execution unit 133, etc.) may be executed by the terminal device 10. Thus, various processes performed by the information processing device 100 may be distributed and processed by a plurality of devices included in the information processing system 1.
[0072] Although the embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and its equivalent scope. Also, these embodiments and their modifications can be appropriately combined as long as the processing contents do not conflict.
Explanation of Reference Numerals
[0073] 1 Information processing system 10 Terminal device 11 Communication unit 12 Input unit 13 Output unit 14 Control unit 100 Information processing device 110 Communication unit 120 Storage unit 121 Assumed answer storage unit 122 VQA model storage unit 130 Control unit 131 Acquisition unit 132 Determination unit 133 Execution unit 134 Provision unit 141 Reception unit 142 Transmission unit N Network
Claims
1. A determination unit that determines whether or not the imaging status of a detection target satisfies a predetermined condition; An execution unit that executes a process using VQA (Visual Question Answering) when it is determined by the determination unit that the predetermined condition is satisfied; An information processing apparatus comprising the same.
2. The determination unit determines whether or not the imaging status satisfies the predetermined condition using the captured image of the detection target. The information processing apparatus according to claim 1, characterized in that.
3. The determination unit determines whether or not the imaging status satisfies the predetermined condition using the captured image captured from a position at a predetermined height. The information processing apparatus according to claim 2, characterized in that.
4. The determination unit determines whether or not the imaging status satisfies the predetermined condition based on a threshold value determined based on the height and a coefficient corresponding to the imaging device. The information processing apparatus according to claim 3, characterized in that.
5. The determination unit determines whether or not the imaging status of the detection target satisfies the predetermined condition by determining whether or not the coordinate data of the detection target satisfies the predetermined condition. The information processing apparatus according to claim 1, characterized in that.
6. The determination unit determines whether or not all the coordinate data of the vertices of a figure having a predetermined shape indicating the detection target satisfies the predetermined condition. The information processing apparatus according to claim 5, characterized in that.
7. The determination unit By determining whether or not the coordinate data exceeds a predetermined threshold value, it is determined whether or not the imaging situation of the detection target satisfies the predetermined conditions The information processing apparatus according to claim 5, characterized in that
8. The determination unit By determining whether or not the range corresponding to the question sentence of VQA satisfies a predetermined condition, it is determined whether or not the imaging situation of the detection target satisfies the predetermined condition The information processing apparatus according to claim 1, characterized in that
9. An information processing method executed by a computer, comprising A determination step of determining whether or not the imaging situation of the detection target satisfies a predetermined condition An execution step of executing a process using VQA (Visual Question Answering) when it is determined by the determination step that the predetermined condition is satisfied An information processing method characterized by including
Citation Information
Patent Citations
State determination device and image analysis device
JP2022071675A