Object detection device, object detection system, object detection method, and recording medium

By compressing and encoding images to generate features for object detection, the object detection device addresses the challenge of separate processing loads, enabling efficient image transmission and processing.

JP7798370B2Active Publication Date: 2026-01-14NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023512574
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-07
Publication Date
2026-01-14
Estimated Expiration
2041-04-07

AI Technical Summary

Technical Problem

Existing object detection devices face challenges in efficiently performing image compression and detection, especially when they are installed in mobile terminals with low processing power, they may transmit the image to an information processing device, they may not have the processing power to process the image to the information processing device, they may not have the processing capacity to perform the compression operation and the compression operation separately and independently.

Method used

The object detection device compresses and encodes images to generate features that enable object detection, reducing the need for separate processing of compression and detection operations.

Benefits of technology

This approach reduces the processing load required for image compression and detection, allowing efficient transmission and processing of images with limited bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798370000001
    Figure 0007798370000001
  • Figure 0007798370000002
    Figure 0007798370000002
  • Figure 0007798370000003
    Figure 0007798370000003
Patent Text Reader

Abstract

An object detection device (1) comprises: a generation means (111) that performs compression-encoding of a first image (IMG_original) which has been acquired from an image generation device and a second image (IMG_target) which indicates a detection target object, in a manner such that feature amounts enabling object detection are extracted and decoding can be performed thereafter, and that thus generates first encoded information (EI_original) which is the compression-encoded first image and which can be used as a first feature amount (CM_original) of the first image and second encoded information (EI_target) which is the compression-encoded second image and which can be used as a second feature amount (CM_target) of the second image; and a detection means (112) that uses the first and second feature amounts to detect the detection target object in the first image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to, for example, the technical fields of an object detection device, an object detection system, an object detection method, and a recording medium that are capable of detecting a detection target object in an image. [Background technology]

[0002] Patent Document 1 describes an example of an object detection device that uses a neural network to detect a detection target object in an image.

[0003] Other prior art documents related to this disclosure include Patent Documents 2 to 4. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2020 / 031422 Brochure [Patent Document 2] Japanese Patent Publication No. 2020-051982 [Patent Document 3] Patent No. 6605742 [Patent Document 4] International Publication No. 2017 / 187516 Brochure Summary of the Invention [Problem to be solved by the invention]

[0005] The object detection device may transmit the image to an external information processing device via a communication line in parallel with the object detection process for detecting a detection target object in the image. For example, if the object detection device is installed in a mobile terminal having a relatively low processing power, the object detection device may transmit the image to an information processing device that can perform information processing on the image that requires a relatively high processing power.

[0006] In this case, to satisfy the bandwidth constraints of the communication line, the object detection device may compress the image and transmit the compressed image to the information processing device. In this case, the object detection device needs to perform a compression operation to compress the image separately and independently from the object detection operation. However, the object detection device does not necessarily have a high enough processing capacity to perform the object detection operation and the compression operation separately and independently. Therefore, it is desirable to reduce the processing load required to perform the object detection operation and the compression operation.

[0007] An object of the present disclosure is to provide an object detection device, an object detection system, an object detection method, and a recording medium that can solve the above-mentioned technical problems. As an example, an object of the present disclosure is to provide an object detection device, an object detection system, an object detection method, and a recording medium that can compress an image and reduce the processing load for detecting a detection target object within the image. [Means for solving the problem]

[0008] The object detection device disclosed herein comprises a generation means that generates first encoded information for the compressed and encoded first image that can be used as a first feature that is the feature of the first image and a second image that shows the object to be detected, by compressing and encoding the first image acquired from an image generation device so as to extract features that enable object detection and so that the images can be later decoded, and second encoded information for the compressed and encoded second image that can be used as a second feature that is the feature of the second image, and a detection means that detects the object to be detected in the first image using the first and second features.

[0009] The object detection system disclosed herein is an object detection system comprising an object detection device and an information processing device, wherein the object detection device comprises a generation means that compresses and encodes a first image acquired from an image generation device and a second image showing a detection target object so as to extract features that enable object detection and so that the images can be later decoded, thereby generating first encoded information that is the compressed and encoded first image and that can be used as a first feature, which is the feature of the first image, and second encoded information that is the compressed and encoded second image and that can be used as a second feature, which is the feature of the second image; a detection means that detects the detection target object in the first image using the first and second features; and a transmission means that transmits the first encoded information to the information processing device via a communication line, and the information processing device performs a predetermined operation using the first encoded information.

[0010] The object detection method disclosed herein compresses and encodes a first image acquired from an image generating device and a second image showing a target object to be detected so as to extract features that enable object detection and so that the images can be later decoded, thereby generating first encoded information that is the compressed and encoded first image and can be used as a first feature that is the feature of the first image, and second encoded information that is the compressed and encoded second image and can be used as a second feature that is the feature of the second image, and detecting the target object in the first image using the first and second features.

[0011] The recording medium disclosed herein is a recording medium having recorded thereon a computer program that causes a computer to execute an object detection method that detects the object to be detected in the first image using the first and second features by compressing and encoding each of a first image acquired from an image generating device and a second image showing the object to be detected so as to extract features that enable object detection and that can be later decoded, thereby generating first encoded information that is the compressed and encoded first image and can be used as a first feature that is the feature of the first image, and second encoded information that is the compressed and encoded second image and can be used as a second feature that is the feature of the second image, and [Effects of the Invention]

[0012] According to the object detection device, object detection system, object detection method, and recording medium described above, it is possible to reduce the processing load for compressing the first image and detecting the detection target object in the first image. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of an object detection system according to this embodiment. [Figure 2] FIG. 2 is a block diagram showing the configuration of the object detection device of this embodiment. [Figure 3] FIG. 3 is a schematic diagram showing the structure of a neural network used by the object detection device of this embodiment. [Figure 4] FIG. 4 is a block diagram showing the configuration of the information processing device of this embodiment. [Figure 5] FIG. 5 is a flowchart showing the flow of operations of the object detection system of this embodiment. [Figure 6] FIG. 6 conceptually illustrates machine learning for generating a computational model used by an object detection device. [Figure 7] FIG. 7 is a schematic diagram showing the structure of a neural network used by an object detection device of the comparative example. [Figure 8]FIG. 8 is a block diagram showing the configuration of an object detection device according to a modified example. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments of an object detection device, an object detection system, an object detection method, and a recording medium will be described with reference to the drawings. Hereinafter, embodiments of an object detection device, an object detection system, an object detection method, and a recording medium will be described using an object detection system SYS to which the embodiments of the object detection device, the object detection system, the object detection method, and the recording medium are applied. However, the present invention is not limited to the embodiments described below.

[0015] <1> Configuration of the object detection system SYS First, the configuration of the object detection system SYS of this embodiment will be described.

[0016] <1-1> Overall configuration of the object detection system SYS First, the overall configuration of the object detection system SYS of this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the overall configuration of the object detection system SYS of this embodiment.

[0017] As shown in Fig. 1, the object detection system SYS includes an object detection device 1 and an information processing device 2. The object detection device 1 and the information processing device 2 can communicate with each other via a communication line 3. The communication line 3 may include a wired communication line. The communication line 3 may include a wired communication line.

[0018] The object detection device 1 is capable of detecting a detection target object in the original image IMG_original. That is, the object detection device 1 is capable of performing object detection. The original image IMG_original is an image in which a detection target object is to be detected. The object detection device 1 may acquire such an original image IMG_original from an image generating device such as a camera. In this embodiment, the object detection device 1 uses a detection target image IMG_target indicating the detection target object to detect the detection target object in the original image IMG_original. That is, the object detection device 1 detects the detection target object indicated by the detection target image IMG_target in the original image IMG_original using the original image IMG_original and the detection target image IMG_target. Specifically, the object detection device 1 generates a feature CM_original of the original image IMG_original based on the original image IMG_original as a feature that enables object detection. Furthermore, based on the detection target image IMG_target, the object detection device 1 generates a feature CM_target of the detection target image IMG_target as a feature that enables object detection. Thereafter, the object detection device 1 detects the detection target object in the original image IMG_original based on the feature CM_original and the feature CM_target.

[0019] The object detection device 1 further compresses and encodes the original image IMG_original so that it can be decoded later. In other words, the object detection device 1 performs a desired compression-encoding process on the original image IMG_original, thereby converting it into a data structure (information format, information form) that can later be decoded corresponding to the desired compression-encoding. Hereinafter, in this application, performing a desired compression-encoding process on an input image having a certain data structure to convert it into a data structure (information format, information form) that can later be decoded corresponding to the desired compression-encoding will be expressed as "compression-encoding the input image so that it can be decoded later" or "compression-encoding the input image so that it can be decoded later." Herein, the term "input image" will be replaced with an image with an appropriate name depending on the description.

[0020] As a result of this compression encoding, the object detection device 1 generates encoded information EI_original, which is the compression-encoded original image IMG_original. The object detection device 1 transmits the generated encoded information EI_original to the information processing device 2 via the communication line 3. As a result, compared to when the original image IMG_original is transmitted to the information processing device 2 via the communication line 3, there is a higher possibility that the bandwidth constraints of the communication line 3 will be satisfied.

[0021] Particularly in this embodiment, the object detection device 1 uses the encoded information EI_original as a feature CM_original of the original image IMG_original (i.e., a feature CM_original for detecting a detection target object). That is, the object detection device 1 generates encoded information EI_original that can be used as the feature CM_original by compression-encoding the original image IMG_original. More specifically, the object detection device 1 generates encoded information EI_original that can be used as the feature CM_original by compression-encoding the original image IMG_original so as to extract a feature that enables object detection and so as to be decodable later (in other words, generates a feature CM_original that can be used as the encoded information EI_original).

[0022] As described above, the object detection device 1 uses the detection target image IMG_target in addition to the original image IMG_original to detect a detection target object. Therefore, in addition to the feature CM_original, the object detection device 1 generates encoded information EI_target, which is the compression-encoded detection target image IMG_target, as the feature CM_target of the detection target image IMG_target (i.e., the feature CM_target for detecting a detection target object). That is, the object detection device 1 generates encoded information EI_target that can be used as the feature CM_target by compressing and encoding the detection target image IMG_target in the same manner as when compressing and encoding the original image IMG_original. More specifically, the object detection device 1 generates encoded information EI_target that can be used as the feature CM_target by compressing and encoding the detection target image IMG_target so as to extract a feature that enables object detection and so that the feature CM_target can be later decoded (in other words, generates the feature CM_target that can be used as the encoded information EI_target). The object detection device 1 may or may not transmit the generated encoded information EI_target to the information processing device 2 via the communication line 3.

[0023] The information processing device 2 receives (i.e., acquires) the encoded information EI_original from the object detection device 1 via the communication line 3. The information processing device 2 performs a predetermined operation using the received encoded information EI_original. In this embodiment, an example will be described in which the information processing device 2 performs a decoding operation to generate a restored image IMG_dec by decoding the encoded information EI_original, as an example of the predetermined operation.

[0024] A specific example of such an object detection system SYS is an augmented reality (AR) system. Augmented reality is a technology that detects a real object present in a real space and places a virtual object at the location of the real object in an image representing the real space. In an augmented reality system, the object detection device 1 may be applied to a mobile terminal such as a smartphone. In this case, the object detection device 1 may detect a detection target object (i.e., a real object) in an original image IMG_original generated by capturing an image of the real space with a camera of the mobile terminal, and place a virtual object at the location of the detected detection target object in the original image IMG_original. In this case, the information processing device 2 may generate a restored image IMG_dec by performing the above-mentioned decoding operation, and further perform an image analysis operation to analyze the restored image IMG_dec. The result of the image analysis operation may be transmitted to the mobile terminal. In this case, the mobile terminal may place the virtual object based on the result of the image analysis operation by the information processing device 2 in addition to the detection result of the detection target object by the object detection device 1. An example of the image analysis operation by the information processing device 2 is an operation of estimating the orientation of the mobile terminal based on the restored image IMG_dec. In this case, the mobile terminal may place a virtual object based on the orientation of the mobile terminal estimated by the image analysis operation by the information processing device 2.

[0025] <1-2> Configuration of object detection device 1 Next, the configuration of the object detection device 1 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the configuration of the object detection device 1.

[0026] 2, the object detection device 1 includes a calculation device 11, a storage device 12, and a communication device 13. The object detection device 1 may further include an input device 14 and an output device 15. However, the object detection device 1 does not necessarily have to include at least one of the input device 14 and the output device 15. The calculation device 11, the storage device 12, the communication device 13, the input device 14, and the output device 15 may be connected via a data bus 16.

[0027] The arithmetic device 11 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic device 11 loads a computer program. For example, the arithmetic device 11 may load a computer program stored in the storage device 12. For example, the arithmetic device 11 may load a computer program stored in a computer-readable, non-transitory storage medium using a storage medium reading device (not shown) included in the object detection device 1. The arithmetic device 11 may acquire (i.e., download or load) the computer program from a device (not shown) located outside the object detection device 1 via the communication device 13 (or another communication device). The arithmetic device 11 executes the loaded computer program. As a result, logical functional blocks for executing operations (in other words, processing) to be performed by the object detection device 1 are realized within the arithmetic device 11. In other words, the arithmetic device 11 can function as a controller for realizing logical functional blocks for executing operations to be performed by the object detection device 1.

[0028] Fig. 2 shows an example of logical functional blocks realized in the arithmetic device 11. As shown in Fig. 2, in the arithmetic device 11, an encoding unit 111 which is a specific example of "generation means", an object detection unit 112 which is a specific example of "detection means", and a transmission control unit 113 which is a specific example of "transmission means" are realized.

[0029] The encoding unit 111 generates encoded information EI_original that can be used as a feature amount CM_original of the original image IMG_original by compression-encoding the original image IMG_original so that it can be decoded later. Furthermore, the encoding unit 111 generates encoded information EI_target that can be used as a feature amount CM_target of the detection target image IMG_target by compression-encoding the detection target image IMG_target so that it can be decoded later.

[0030] The object detection unit 112 detects a detection target object in the original image IMG_original based on the feature amount CM_origin and the feature amount CM_target generated by the encoding unit 111.

[0031] In this embodiment, the encoding unit 111 generates encoded information EI_original and EI_target (i.e., feature quantities CM_original and CM_target) using a computation model generated by machine learning. Furthermore, the object detection unit 112 detects a detection target object in the original image IMG_original using a computation model generated by machine learning.

[0032] The computational model may include a compression encoding model and an object detection model. The compression encoding model may be a model for mainly generating encoded information EI_original and EI_target (i.e., feature quantities CM_origin and CM_target). The object detection model may be a model for mainly detecting a detection target object in the original image IMG_original based on the feature quantities CM_origin and CM_target (i.e., encoded information EI_original and EI_target).

[0033] An example of a computational model generated by machine learning is a neural network NN. An example of the neural network NN used by the encoding unit 111 and the object detection unit 112 is shown schematically in Fig. 3. As shown in Fig. 3, the neural network NN includes a network portion NN1 which is a specific example of a "first model portion" and a network portion NN2 which is a specific example of a "second model portion."

[0034] The network part NN1 is used by the encoding unit 111 mainly to generate encoded information EI_original and EI_target (i.e., feature amounts CM_original and CM_target). In other words, the network part NN1 is a neural network for realizing the above-mentioned compression encoding model. When an input image is input, the network part NN1 can output encoded information that is the input image that has been compression encoded so that it can be decoded later and that can be used as feature amounts of the input image. Therefore, when an original image IMG_original is input to the network part NN1, the network part NN1 outputs encoded information EI_original (i.e., feature amount CM_original). When a detection target image IMG_target is input to the network part NN1, the network part NN1 outputs encoded information EI_target (i.e., feature amount CM_target).

[0035] The network portion NN1 may include a neural network conforming to a desired compression encoding method. For example, an encoder portion of an autoencoder may be used as the network portion NN1. In this case, the information processing device 2 may generate the restored image IMG_dec from the encoded information EI_original using a decoder portion of the autoencoder.

[0036] The network part NN2 is used by the object detection unit 112 mainly to detect a detection target object in the original image IMG_original. In other words, the network part NN2 is a neural network for realizing the above-mentioned object detection model. When the feature of one image and the feature of another image are input, the network part NN2 outputs a detection result of an object indicated by the other image in the one image. The feature CM_original and CM_target output from the network part NN1 are input to the network part NN2. In this case, the network part NN2 outputs a detection result of a detection target object indicated by the detection target image IMG_target in the original image IMG_original. For example, the network part NN2 may output information regarding the presence or absence of the detection target object in the original image IMG_original as the detection result of the detection target object. The network part NN2 may also output information regarding the position (e.g., the position of the bounding box) of the detection target image IMG_target in the original image IMG_original as the detection result of the detection target object.

[0037] The network portion NN2 may include a neural network conforming to a desired object detection method for detecting an object using two images, such as SiamRPN (Siamese Region Proposal Network).

[0038] 2 again, the transmission control unit 113 uses the communication device 13 to transmit the encoded information EI_original generated by the encoding unit 111 to the information processing device 2. More specifically, as shown in FIG. 3, the transmission control unit 113 uses the communication device 13 to transmit the encoded information EI_original output by the network part NN1 to the information processing device 2. Furthermore, the transmission control unit 113 may use the communication device 13 to transmit the encoded information EI_target generated by the encoding unit 111 to the information processing device 2. More specifically, as shown in FIG. 3, the transmission control unit 113 may use the communication device 13 to transmit the encoded information EI_target output by the network part NN1 to the information processing device 2.

[0039] The storage device 12 can store desired data. For example, the storage device 12 may temporarily store a computer program executed by the arithmetic device 11. The storage device 12 may temporarily store data that the arithmetic device 11 temporarily uses when the arithmetic device 11 is executing a computer program. The storage device 12 may store data that the object detection device 1 stores long-term. The storage device 12 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 12 may include a non-temporary recording medium.

[0040] The communication device 13 is capable of communicating with the information processing device 2 via the communication line 3. In this embodiment, the communication device 13 transmits the encoded information EI_original to the information processing device 2 via the communication line 3 under the control of the transmission control unit 113. Furthermore, the communication device 13 may transmit the encoded information EI_target to the information processing device 2 via the communication line 3 under the control of the transmission control unit 113.

[0041] Input device 14 is a device that accepts information input to object detection device 1 from outside object detection device 1. For example, input device 14 may include an operation device (for example, at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of object detection device 1. For example, input device 14 may include a reading device that can read information recorded as data on a recording medium that can be externally attached to object detection device 1.

[0042] The output device 15 is a device that outputs information to the outside of the object detection device 1. For example, the output device 15 may output information as an image. That is, the output device 15 may include a display device (a so-called display) that can display an image showing the information to be output. For example, the output device 15 may output information as sound. That is, the output device 15 may include an audio device (a so-called speaker) that can output sound. For example, the output device 15 may output information on paper. That is, the output device 15 may include a printing device (a so-called printer) that can print desired information on paper.

[0043] <1-3> Configuration of information processing device 2 Next, the configuration of the information processing device 2 will be described with reference to Fig. 4. Fig. 4 is a block diagram showing the configuration of the information processing device 2.

[0044] 4, the information processing device 2 includes a calculation device 21, a storage device 22, and a communication device 23. The information processing device 2 may further include an input device 24 and an output device 25. However, the information processing device 2 does not necessarily have to include at least one of the input device 24 and the output device 25. The calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.

[0045] The arithmetic device 21 includes, for example, at least one of a CPU, a GPU, and an FPGA. The arithmetic device 21 reads a computer program. For example, the arithmetic device 21 may read a computer program stored in the storage device 22. For example, the arithmetic device 21 may read a computer program stored in a computer-readable, non-transitory storage medium using a storage medium reading device (not shown) included in the information processing device 2. The arithmetic device 21 may acquire (i.e., download or read) the computer program from a device (not shown) located outside the information processing device 2 via the communication device 23 (or another communication device). The arithmetic device 21 executes the read computer program. As a result, logical functional blocks for executing operations to be performed by the information processing device 2 are realized within the arithmetic device 21. In other words, the arithmetic device 21 can function as a controller for realizing logical functional blocks for executing operations to be performed by the information processing device 2.

[0046] FIG. 4 shows an example of logical functional blocks realized in the arithmetic device 21. As shown in FIG. 4, an information acquisition unit 211 and a processing unit 212 are realized in the arithmetic device 21. The information acquisition unit 211 receives (i.e., acquires) the encoded information EI_original transmitted from the object detection device 1 using the communication device 23. The processing unit 212 performs a predetermined operation using the encoded information EI_original. In this embodiment, the processing unit 212 performs a decoding operation to generate a restored image IMG_dec by decoding the encoded information EI_original acquired by the information acquisition unit 211. Furthermore, the processing unit 212 may perform an image analysis operation to analyze the restored image IMG_dec.

[0047] The storage device 22 can store desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic device 21. The storage device 22 may temporarily store data that the arithmetic device 21 temporarily uses when the arithmetic device 21 is executing a computer program. The storage device 22 may store data that the information processing device 2 stores for a long period of time. The storage device 22 may include at least one of a RAM, a ROM, a hard disk device, a magneto-optical disk device, an SSD, and a disk array device. In other words, the storage device 22 may include a non-temporary recording medium.

[0048] The communication device 23 is capable of communicating with the object detection device 1 via the communication line 3. In this embodiment, the communication device 23 may receive (i.e., acquire) the encoded information EI_original from the object detection device 1 via the communication line 3 under the control of the information acquisition unit 211.

[0049] The input device 24 is a device that accepts information input to the information processing device 2 from outside the information processing device 2. For example, the input device 24 may include an operation device (for example, at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the information processing device 2. For example, the input device 24 may include a reading device that can read information recorded as data on a recording medium that can be externally attached to the information processing device 2.

[0050] The output device 25 is a device that outputs information to the outside of the information processing device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (a so-called display) that can display an image showing the information to be output. For example, the output device 25 may output information as sound. That is, the output device 25 may include an audio device (a so-called speaker) that can output sound. For example, the output device 25 may output information on paper. That is, the output device 25 may include a printing device (a so-called printer) that can print desired information on paper.

[0051] <2> Operation of the object detection system SYS Next, the operation performed by the object detection system SYS will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of the operation performed by the object detection system SYS.

[0052] As shown in FIG. 5, the object detection device 1 (particularly, the encoding unit 111) acquires an original image IMG_original (step S11). For example, the object detection device 1 may acquire the original image IMG_original from a camera, which is a specific example of an image generating device. In this case, the object detection device 1 may acquire the original image IMG_original from the camera every time the camera generates an original image IMG_original. The object detection device 1 may acquire multiple original images IMG_original as time-series data from the camera. In this case, the operation shown in FIG. 5 is performed using each original image IMG_.

[0053] Furthermore, the object detection device 1 (particularly, the encoding unit 111) acquires the detection target image IMG_target (step S11). For example, if the detection target image IMG_target is stored in the storage device 12, the object detection device 1 may acquire the detection target image IMG_target from the storage device 12. For example, if the detection target image IMG_target is recorded on a recording medium that can be attached externally to the object detection device 1, the object detection device 1 may acquire the detection target image IMG_target from the recording medium using a recording medium reading device (for example, the input device 14) provided in the object detection device 1. For example, if the detection target image IMG_target is recorded on a device (for example, a server) external to the object detection device 1, the object detection device 1 may acquire the detection target image IMG_target from the external device using the communication device 13.

[0054] Note that if the detection target object does not change, the object detection device 1 does not need to acquire the detection target image IMG_target again after once acquiring the detection target image IMG_target. In other words, the object detection device 1 may acquire the detection target image IMG_target when the detection target object changes.

[0055] Thereafter, the object detection device 1 (particularly, the encoding unit 111) generates encoded information EI_original that can be used as a feature amount CM_original of the original image IMG_original by compression-encoding the original image IMG_original so that it can be decodable later (step S12). Furthermore, the object detection device 1 (particularly, the encoding unit 111) generates encoded information EI_target that can be used as a feature amount CM_target of the detection target image IMG_target by compression-encoding the detection target image IMG_target so that it can be decodable later (step S12).

[0056] Thereafter, the object detection device 1 (particularly the object detection unit 112) detects the detection target object in the original image IMG_original based on the feature amounts CM_original and CM_target generated in step S12 (step S13). The operation of detecting the detection target object may include an operation of detecting a region of a desired shape (e.g., a rectangular region, a so-called bounding box) that includes the detection target object in the original image IMG_original. The operation of detecting the detection target object may include an operation of detecting the position (e.g., coordinate values) of the region of a desired shape that includes the detection target object in the original image IMG_original. The operation of detecting the detection target object may include an operation of detecting characteristics of the detection target object in the original image IMG_original (e.g., at least one of color, shape, size, and orientation).

[0057] The detection result of the detection target object in step S13 may be used for a desired purpose. For example, as described above, the detection result of the detection target object in step S13 may be used for AR purposes. That is, the detection result of the detection target object in step S13 may be used for placing a virtual object at the position of the detection target object.

[0058] In parallel with or before or after the operation of step S13, the object detection device 1 (particularly, the transmission control unit 113) uses the communication device 13 to transmit the encoded information EI_original generated in step S12 to the information processing device 2 (step S14). Here, because the encoded information EI_original is the compression-encoded original image IMG_original, the data size of the encoded information EI_original is smaller than the data size of the original image IMG_original. Therefore, compared to when the original image IMG_original is transmitted to the information processing device 2 via the communication line 3, there is a higher possibility that the bandwidth constraints of the communication line 3 will be satisfied. In other words, even when the bandwidth of the communication line 3 is relatively narrow (i.e., the amount of data that can be transmitted per unit time is relatively small), the object detection device 1 can transmit the encoded information EI_original to the information processing device 2.

[0059] As a result, the information processing device 2 (particularly, the information acquisition unit 211) receives the encoded information EI_original transmitted from the object detection device 1 using the communication device 23 (step S21). Thereafter, the information processing device 2 (particularly, the processing unit 212) performs a predetermined operation using the encoded information EI_original (step S22). For example, the processing unit 212 may perform a decoding operation to generate a restored image IMG_dec by decoding the encoded information EI_original acquired by the information acquisition unit 211. The processing unit 212 may also perform an image analysis operation to analyze the restored image IMG_dec.

[0060] <3> Generating computational models using machine learning Next, machine learning for generating a computational model used by the object detection device 1 will be described with reference to Fig. 6. Fig. 6 conceptually illustrates machine learning for generating a computational model used by the object detection device 1. For convenience of explanation, the following description will focus on machine learning performed when the computational model is the neural network NN of Fig. 3. However, even if the computational model is different from the neural network NN of Fig. 3, the computational model may be generated by the machine learning described below.

[0061] The neural network NN is generated by machine learning using a training dataset including a plurality of training data in which a training image (hereinafter referred to as a "training image IMG_learn_original") and a correct label y_learn of the detection result of the detection target object in the training image IMG_learn_original are associated with each other. Furthermore, even after the neural network NN is once generated, the neural network NN may be updated as appropriate by machine learning using a training dataset including new training data.

[0062] To generate or update a neural network NN, a training image IMG_learn_original included in the training data is input to a network portion NN1 (i.e., a compression-encoded model) included in an initial or generated neural network NN. As a result, the network portion NN1 compression-encodes the training image IMG_learn_original so that it can be later decodable, and outputs encoded information EI_learn, which is the compressed and encoded training image IMG_learn_original and can be used as a feature CM_learn_original of the training image IMG_learn_original. Furthermore, a detection target image for training (hereinafter referred to as a "detection target image IMG_learn_target") indicating a detection target object for training is input to the network portion NN1 included in the initial or generated neural network NN. As a result, the network portion NN1 compression-encodes the detection target image IMG_learn_target so that it can be later decodable, and outputs encoded information EI_learn_target, which is the compressed and encoded detection target image IMG_learn_target and can be used as a feature CM_learn_target of the detection target image IMG_learn_target.

[0063] Thereafter, the output of the network part NN1 (i.e., the features CM_learn_original and CM_learn_target) is input to the network part NN2 (i.e., the object detection model) included in the initial or generated neural network NN. As a result, the network part NN2 outputs the actual detection result y of the detection target object in the learning image IMG_learn_original. Furthermore, the encoded information EI_learn_original output by the network part NN2 is decoded. As a result, the restored image IMG_learn_dec is generated.

[0064] The above operations are repeated for multiple pieces of training data (or a part of them) included in the training data set. Furthermore, the operations performed for multiple pieces of training data (or a part of them) may be repeated for multiple detection target images IMG_learn_target.

[0065] Thereafter, a neural network NN is generated or updated using a loss function Loss, which includes a loss function Loss1 related to detection of the detection target object and a loss function Loss2 related to compression encoding and decoding. The loss function Loss1 is a loss function related to the error between the output y of the network part NN2 (i.e., the actual detection result of the detection target object in the training image IMG_learn_original by the network part NN2) and the correct label y_learn. For example, the loss function Loss1 may be a loss function that decreases as the error between the output y of the network part NN2 and the correct label y_learn decreases. On the other hand, the loss function Loss2 is a loss function related to the error between the restored image IMG_learn_dec and the training image IMG_learn_original. For example, the loss function Loss2 may be a loss function that decreases as the error between the restored image IMG_learn_dec and the training image IMG_learn_original decreases.

[0066] The neural network NN may be generated or updated so that the loss function Loss is minimized. In this case, the neural network NN may be generated or updated using an existing algorithm for machine learning so that the loss function Loss is minimized. For example, the neural network NN may be generated or updated using the backpropagation method so that the loss function Loss is minimized. As a result, the neural network NN is generated or updated.

[0067] <4> Technical effects of the object detection system SYS As described above, in this embodiment, the object detection device 1 generates encoded information EI_original that can be used as feature quantities CM_original of the original image IMG_original by compression-encoding the original image IMG_original. That is, the object detection device 1 does not need to perform the operation of generating the feature quantities CM_original and the operation of generating the encoded information EI_original separately and independently. The object detection device 1 does not need to perform the operation of generating the feature quantities CM_original separately and independently from the encoded information EI_original. The object detection device 1 does not need to perform the operation of generating the encoded information EI_original separately and independently from the feature quantities CM_original. This makes it possible to reduce the processing load required to compress the original image IMG_original and detect the detection target object in the original image IMG_original.

[0068] Specifically, the comparative object detection device, which does not generate encoded information EI_original that can be used as the feature CM_original, must perform the operation of generating the feature CM_original and the operation of generating the encoded information EI_original separately and independently, as shown in FIG. 7 . In the example shown in FIG. 7 , the comparative object detection device compresses the original image IMG_original and detects the target object in the original image IMG_original using a neural network NN′ including a network portion NN3 for generating the feature CM_original separately from the encoded information EI_original, a network portion NN4 for generating the encoded information EI_original separately from the feature CM_original, and a network portion NN2 for detecting the target object based on the feature CM_original. Unlike the comparative object detection device, the object detection device 1 of this embodiment does not need to include either the network portion NN3 or NN4. Therefore, the structure of the neural network NN used by the object detection device 1 is simpler than the structure of the neural network NN′ used by the comparative object detection device. That is, the structure of the computation model used by the object detection device 1 is simpler than the structure of the computation model used by the object detection device of the comparative example. As a result, in this embodiment, the processing load for compressing the original image IMG_original and generating the feature quantity CM_original of the original image IMG_original can be reduced compared to the comparative example. That is, in this embodiment, the processing load for compressing the original image IMG_original and detecting the detection target object in the original image IMG_original can be reduced compared to the comparative example.

[0069] Furthermore, the neural network NN (i.e., the computational model) is generated by machine learning using a loss function Loss that includes a loss function Loss1 related to detection of the detection target object and a loss function Loss2 related to compression encoding and decoding. Therefore, a computational model is generated that can appropriately generate encoded information EI_original that is the compressed original image IMG_original and that can be used as the feature quantity CM_original of the original image IMG_original. As a result, the object detection device 1 can appropriately generate encoded information EI_original that is the compressed original image IMG_original and that can be used as the feature quantity CM_original of the original image IMG_original by compression encoding the original image IMG_original using the computational model generated in this manner.

[0070] <5> Variations In the above description, the object detection device 1 transmits the encoded information EI_original to the information processing device 2. However, the object detection device 1 does not have to transmit the encoded information EI_original to the information processing device 2. For example, the object detection device 1 may store the encoded information EI_original in the storage device 12. In this case, as shown in FIG. 8 , the object detection device 1 does not have to include the transmission control unit 113.

[0071] In the above description, the object detection device 1 detects the detection target object indicated by the detection target image IMG_target in the original image IMG_original using the original image IMG_original and the detection target image IMG_target. However, the object detection device 1 may detect the detection target object in the original image IMG_original without using the detection target image IMG_target. For example, the object detection device 1 may detect the target object using a computational model conforming to a desired object detection method for detecting the object using the image in which the object is to be detected. An example of a computational model conforming to a desired object detection method for detecting the object using the image in which the object is to be detected is a computational model conforming to YOLO (You Only Look Once). Even in this case, the object detection device 1 may generate encoded information EI_original, which is the compressed and encoded original image IMG_original and can be used as the feature quantity CM_original of the original image IMG_original, by compression-encoding the original image IMG_original so that it can be later decoded. As a result, the object detection device 1 can enjoy the above-mentioned effects.

[0072] As an example, when the above-described YOLO-compliant computational model is used, machine learning of the YOLO-compliant computational model may be performed so that the output of the intermediate layer of the YOLO-compliant computational model is decodable. In other words, machine learning may be performed to generate a computational model that is compliant with YOLO but extends YOLO so as to include an intermediate layer whose output is later decodable. As a result, the intermediate layer of the YOLO-compliant computational model can output coded information that can be used as a feature for object detection and can later be decodable. Therefore, even an object detection device 1 that performs object detection using a YOLO-compliant computational model can enjoy the above-described effects.

[0073] In the above description, the information processing device 2 performs, as examples of predetermined operations, a decoding operation of generating a restored image IMG_dec by decoding the coded information EI_original, and an image analysis operation of analyzing the restored image IMG_dec. However, the information processing device 2 may perform operations other than the decoding operation and the image analysis operation. For example, the information processing device 2 may perform an operation of storing the coded information EI_original received from the object detection device 1 in the storage device 22. For example, the information processing device 2 may perform an operation of storing the restored image IMG_dec generated from the coded information EI_original in the storage device 22.

[0074] <6> Additional notes The following additional notes are provided regarding the above-described embodiment. [Appendix 1] a generation means for generating first encoded information that is the compressed and encoded first image and that can be used as a first feature amount that is the feature amount of the first image, and second encoded information that is the compressed and encoded second image and that can be used as a second feature amount that is the feature amount of the second image, by compressing and encoding each of a first image acquired from an image generation device and a second image showing a detection target object so as to extract a feature amount that enables object detection and so as to be decodable later; a detection means for detecting the detection target object in the first image using the first and second feature amounts; An object detection device comprising: [Appendix 2] The information processing device further includes a transmitting means for transmitting the first encoded information via a communication line to an information processing device that performs a predetermined operation using the first encoded information. 2. The object detection device of claim 1. [Appendix 3] The predetermined operation includes at least one of a first operation of generating a third image by decoding the first encoded information, a second operation of analyzing the third image, a third operation of storing the first encoded information in a storage device, and a fourth operation of storing the third image in a storage device. 3. The object detection device according to claim 2. [Appendix 4] the generating means generates the first and second encoded information usable as the first and second feature amounts, respectively, by using a first model portion of a computational model generated by machine learning that outputs the first and second encoded information when the first and second images are input; the detection means detects the detection target object using a second model portion of the computational model that outputs a detection result of the detection target object in the first image when the first and second feature amounts are input; The computational model is generated by machine learning using a first loss function based on an error between the detection result of the detection target object output by a second model portion of the computational model to which a fourth image for learning has been input and a correct label of the detection result of the detection target object in the fourth image, and a second loss function based on an error between the fourth image and a third image generated by decoding the first encoded information output by the first model portion of the computational model to which the fourth image has been input. 4. An object detection device according to any one of claims 1 to 3. [Appendix 5] the computational model includes a neural network; the first model portion includes an encoder portion of an autoencoder; 5. The object detection device of claim 4. [Appendix 6] An object detection system including an object detection device and an information processing device, The object detection device a generation means for generating first encoded information that is the compressed and encoded first image and that can be used as a first feature amount that is the feature amount of the first image, and second encoded information that is the compressed and encoded second image and that can be used as a second feature amount that is the feature amount of the second image, by compressing and encoding each of a first image acquired from an image generation device and a second image showing a detection target object so as to extract a feature amount that enables object detection and so as to be decodable later; a detection means for detecting the detection target object in the first image using the first and second feature amounts; a transmitting means for transmitting the first encoded information to the information processing device via a communication line; Equipped with The information processing device performs a predetermined operation using the first encoded information. Object detection system. [Appendix 7] a first image acquired from an image generating device and a second image showing a detection target object are each compression-encoded so as to extract a feature amount that enables object detection and so as to be decodable later, thereby generating first encoded information that is the compression-encoded first image and can be used as a first feature amount that is the feature amount of the first image, and second encoded information that is the compression-encoded second image and can be used as a second feature amount that is the feature amount of the second image; Detecting the detection target object in the first image using the first and second feature amounts Object detection methods. [Appendix 8] On the computer, a first image acquired from an image generating device and a second image showing a detection target object are each compression-encoded so as to extract a feature amount that enables object detection and so as to be decodable later, thereby generating first encoded information that is the compression-encoded first image and can be used as a first feature amount that is the feature amount of the first image, and second encoded information that is the compression-encoded second image and can be used as a second feature amount that is the feature amount of the second image; Detecting the detection target object in the first image using the first and second feature amounts A recording medium on which a computer program for executing an object detection method is recorded.

[0075] At least some of the constituent elements of each of the above-described embodiments can be appropriately combined with at least some of the other constituent elements of each of the above-described embodiments. Some of the constituent elements of each of the above-described embodiments may not be used. Furthermore, to the extent permitted by law, the disclosures of all documents (e.g., published patent applications) cited in this disclosure are incorporated by reference as part of the description of this disclosure.

[0076] This disclosure may be modified as appropriate within the scope of the claims and the technical idea that can be read from the entire specification. Object detection devices, object detection systems, object detection methods, and recording media that incorporate such modifications are also included in the technical idea of ​​this disclosure. [Explanation of symbols]

[0077] SYS Object Detection System 1. Object detection device 11 Arithmetic unit 111 Encoding section 112 Object detection unit 113 Transmission control section 2. Information processing equipment IMG_original Original image IMG_target Detection target image EI_original, EI_target encoding information CM_original, CM_target features NN neural network NN1 and NN2 network parts

Claims

1. a generating means for generating first encoded information that is compressed and usable as a first feature of the first image and second encoded information that is compressed and usable as a second feature of the second image by compressing and encoding each of a first image acquired from an image generating device and a second image showing a detection target object so as to extract a feature that enables object detection and to be decodable later; and a detection means for detecting the detection target object in the first image using the first and second feature amounts; a transmitting means for transmitting the first encoded information via a communication line to an information processing device that performs a predetermined operation using the first encoded information and that is capable of performing information processing on an image that requires a relatively high processing capacity; Equipped with the compression encoding is realized so as to satisfy the bandwidth constraints of the communication line; the predetermined operation includes at least one of a first operation of generating a third image by decoding the first encoded information, a second operation of analyzing the third image, a third operation of storing the first encoded information in a storage device, and a fourth operation of storing the third image in a storage device. An object detection device installed in a terminal device having relatively low processing power.

2. the generating means generates the first and second encoded information usable as the first and second feature amounts, respectively, by using a first model portion of a computational model generated by machine learning that outputs the first and second encoded information when the first and second images are input; the detection means detects the detection target object using a second model portion of the computational model that outputs a detection result of the detection target object in the first image when the first and second feature amounts are input; The computational model is generated by machine learning using a first loss function based on an error between the detection result of the detection target object output by a second model portion of the computational model to which a fourth image for learning has been input and a correct label of the detection result of the detection target object in the fourth image, and a second loss function based on an error between the fourth image and a third image generated by decoding the first encoded information output by the first model portion of the computational model to which the fourth image has been input. The object detection device according to claim 1 .

3. the computational model includes a neural network; the first model portion includes an encoder portion of an autoencoder; The object detection device according to claim 2 .

4. An object detection system comprising an object detection device mounted on a terminal device having a relatively low processing capacity, and an information processing device capable of performing information processing on an image that requires a relatively high processing capacity, The object detection device a generating means for generating first encoded information that is compressed and usable as a first feature of the first image and second encoded information that is compressed and usable as a second feature of the second image by compressing and encoding each of a first image acquired from an image generating device and a second image showing a detection target object so as to extract a feature that enables object detection and to be decodable later; and a detection means for detecting the detection target object in the first image using the first and second feature amounts; a transmitting means for transmitting the first encoded information to the information processing device via a communication line; Equipped with the compression encoding is realized so as to satisfy the bandwidth constraints of the communication line; the information processing device performs a predetermined operation using the first encoded information; the predetermined operation includes at least one of a first operation of generating a third image by decoding the first encoded information, a second operation of analyzing the third image, a third operation of storing the first encoded information in a storage device, and a fourth operation of storing the third image in a storage device. Object detection system.

5. compressing and encoding each of a first image acquired from an image generating device and a second image showing a detection target object so as to extract a feature amount that enables object detection and so as to be decodable later, thereby generating first encoded information that is compressed and usable as a first feature amount of the first image and second encoded information that is compressed and usable as a second feature amount of the second image; Detecting the detection target object in the first image using the first and second feature amounts; transmitting the first encoded information to an information processing device that performs a predetermined operation using the first encoded information and that is capable of performing information processing on an image that requires a relatively high processing capacity; the compression encoding is realized so as to satisfy the bandwidth constraints of the communication line; the predetermined operation includes at least one of a first operation of generating a third image by decoding the first encoded information, a second operation of analyzing the third image, a third operation of storing the first encoded information in a storage device, and a fourth operation of storing the third image in a storage device. An object detection method executed by an object detection device installed in a terminal device having relatively low processing power.

6. A computer that is an object detection device mounted on a terminal device having a relatively low processing power, compressing and encoding each of a first image acquired from an image generating device and a second image showing a detection target object so as to extract a feature amount that enables object detection and so as to be decodable later, thereby generating first encoded information that is compressed and usable as a first feature amount of the first image and second encoded information that is compressed and usable as a second feature amount of the second image; Detecting the detection target object in the first image using the first and second feature amounts; transmitting the first encoded information to an information processing device that performs a predetermined operation using the first encoded information and that is capable of performing information processing on an image that requires a relatively high processing capacity; the compression encoding is realized so as to satisfy the bandwidth constraints of the communication line; the predetermined operation includes at least one of a first operation of generating a third image by decoding the first encoded information, a second operation of analyzing the third image, a third operation of storing the first encoded information in a storage device, and a fourth operation of storing the third image in a storage device. A computer program that causes an object detection method to be performed.

Citation Information

Patent Citations

  • Information processor, store system and program

    JP2013182323A

  • Learning program, learning method and object detecting apparatus

    JP2018205920A

  • Shelf management system and program

    JP2019200697A

  • Image inspection device and inspection model construction system

    JP2020051982A

  • Adaptive artificial neural network selection techniques.

    JP6605742B2