Information processing device and information processing method
By constructing virtual recognition targets in a virtual space from real-space estimates and annotating events, the method addresses the inefficiencies of real-space training data, enhancing the accuracy and detail of image recognition models.
Patent Information
- Application Number
- PCT/JP2024/004146
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2025-08-14
AI Technical Summary
Obtaining training data for image recognition models is time-consuming and costly when using real-space images, and annotations from real space have low accuracy, while virtual space images lack comprehensive reproduction of real-world situations, leading to unnatural scenarios.
Construct a virtual recognition target in a virtual space based on estimated information from real-space images using a first image recognition model, acquire virtual space images, and generate training data by annotating events onto these images to create highly accurate training data for the model.
This method allows for highly detailed and accurate training data generation, improving the accuracy of image recognition models by reproducing real-space scenarios and reducing the need for optical capturing.
Smart Images

Figure JP2024004146_14082025_PF_FP_ABST
Abstract
Description
Information processing device and information processing method
[0001] The present invention relates to an information processing device and an information processing method.
[0002] In order to train an image recognition model to recognize a predetermined event related to a recognition target depicted in an image, training data consisting of pairs of images and annotations indicating the correct answer for the event is required. Also, a technique for generating training data for training an image recognition model using an image of a recognition target configured in a virtual space is known. For example, Patent Document 1 listed below discloses a technique in which a renderer generates a composite image of a scene using a content library or the like, and then uses the composite image to generate training data for a real-world navigation application.
[0003] Japanese Patent Application Laid-Open No. 2021-140767
[0004] Obtaining training data images by capturing images of real space was time-consuming and costly. Furthermore, the accuracy of annotations obtained by capturing events related to the recognition target from images of real space was low. On the other hand, annotations corresponding to images captured by capturing images of a virtual space simulating real space have high accuracy. However, it was difficult to comprehensively reproduce various situations of the recognition target that may occur in real space in a recognition target configured in a virtual space. Furthermore, there were cases where training data images were generated based on a virtual space that reproduced situations that would be unnatural in real space.
[0005] Therefore, an object of the present disclosure is to obtain highly accurate training data for an image recognition model.
[0006] In order to solve the above problem, an information processing device according to one aspect of the present disclosure includes an estimation unit that estimates estimated information including at least one event based on a real-space image using a first image recognition model that recognizes an event related to a recognition target represented in an image; a virtual space construction unit that constructs a virtual recognition target, which is a virtual recognition target that reproduces the event, in a virtual space representing real space based on the estimated information; a virtual space image acquisition unit that acquires a virtual space image, which is an image obtained by projecting a virtual space including the virtual recognition target from a viewpoint position within the virtual space; and a training data construction unit that outputs training data for the image recognition model, which includes a virtual space image associated with an event included in the estimated information.
[0007] According to the above aspect, a virtual recognition object that reproduces an event recognized by an image recognition model is constructed based on estimated information, thereby constructing a virtual space that could occur in real space. Training data is then constructed from virtual space images projected from the virtual space that could occur in real space, thereby obtaining training data suitable for learning an image recognition model to recognize images of real space. Furthermore, since the virtual space images included in the training data are acquired by projecting the virtual space, they can be more highly detailed than images optically captured from the real space. This makes it possible to improve the accuracy of machine learning for the image recognition model. Furthermore, training data is constructed by annotating events based on the estimated information onto virtual space images projected from a virtual space including a virtual recognition object reproduced based on estimated information, thereby making it possible to obtain training data with extremely high accuracy as correct answer information in image recognition.
[0008] It is possible to obtain highly accurate training data for image recognition models.
[0009] FIG. 1 is a block diagram showing the functional configuration of an information processing device of this embodiment. FIG. 2 is a diagram showing an acquisition process for acquiring estimated information from a real space image. FIG. 3 is a diagram showing a schematic diagram of a virtual space in which a person to be recognized is placed based on estimated information. FIG. 4 is a diagram showing the acquisition of a virtual space image captured by a virtual camera installed in the virtual space and the configuration of training data. FIG. 5 is a diagram showing a process for generating an image recognition model by machine learning using training data. FIG. 6 is a flowchart showing the processing contents of an information processing method in an information processing system. FIG. 7 is a diagram showing the configuration of an information processing program. FIG. 8 is a hardware block diagram of the information processing device.
[0010] An information processing device according to an embodiment of the present invention will be described with reference to the drawings. Whenever possible, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0011] 1 is a block diagram showing the device configuration of an information processing system including an information processing device according to this embodiment, and the functional configuration of the information processing device. The information processing system 1 is a system that generates training data for learning an image recognition model that recognizes events related to a recognition target depicted in an image, and generates training data for an image recognition model that recognizes the facial orientation of a person depicted in an image in real space, as an example. The event recognized by the image recognition model is not limited to the facial orientation of a person; for example, a person with bad behavior in an image may be the event to be recognized.
[0012] 1, the information processing device 10 functionally comprises a real space image acquisition unit 11, an estimation unit 12, a virtual space construction unit 13, a virtual space image acquisition unit 14, a training data construction unit 15, and a model generation unit 16. In the example shown in Fig. 1, the functional units 11 to 16 are configured in one information processing device 10, but they may also be configured in a distributed manner across multiple devices.
[0013] Each functional unit of the information processing device 10 is configured to be able to access a storage means (storage) such as a training data storage unit 21. The training data storage unit 21 stores generated training data. In the example shown in Fig. 1, the training data storage unit 21 is configured in another device accessible from the information processing device 10, but it may also be configured in the information processing device 10.
[0014] Next, a description will be given of each functional unit of the information processing device 10. The real space image acquisition unit 11 acquires a real space image obtained by capturing an image of a real space including a recognition target. Specifically, the real space image acquisition unit 11 acquires a real space image obtained by capturing an image of a real space in which the recognition target exists using a camera.
[0015] 2 is a diagram showing a process for acquiring a real-space image and acquiring estimated information from the real-space image. For example, the real-space image acquisition unit 11 acquires a real-space image rp depicting a person hr, which is an optically captured image of a real space rs, such as a store, where the person hr is located. The real-space image rp is an optically captured image and is therefore composed of pixel data.
[0016] The estimation unit 12 estimates estimated information including at least one event based on the real space image using a first image recognition model that recognizes an event related to a predetermined recognition target represented in the image.
[0017] Specifically, the estimation unit 12 uses a first image recognition model md1 that recognizes a predetermined event to recognize the event based on the real space image rp, and outputs estimated information ri that is made up of data indicating the recognized event.
[0018] For example, the first image recognition model md1 is a model that recognizes a person hr as a recognition target and recognizes the facial orientation of the person as a predetermined event. As shown in Fig. 2, the estimation unit 12 uses the first image recognition model md1 to output estimated information ri including data indicating the facial orientation of the person hr based on the real-space image rp. The estimated information ri may include data such as yaw, pitch, and roll that indicate the facial orientation.
[0019] The virtual space construction unit 13 constructs a virtual recognition object, which is a virtual recognition object that reproduces a predetermined event, in a virtual space that represents a real space based on the estimated information ri. Specifically, the virtual space construction unit 13 uses a three-dimensional virtual space model that represents the real space rs to place the virtual recognition object that reproduces the event indicated in the estimated information ri within the virtual space represented by the virtual space model.
[0020] 3 is a diagram schematically illustrating a virtual space in which a person to be recognized is placed based on estimated information. In the example shown in FIG. 3, the virtual space construction unit 13 places a person hm whose facial orientation represented by estimated information ri within a virtual space vs represented by a virtual space model. The position at which the person hm is placed can be obtained by a known analysis process based on a real-space image rp.
[0021] The three-dimensional virtual space model may be generated by any known method. For example, the three-dimensional virtual space model may be generated based on a two-dimensional real-space image. That is, the depth of each pixel in the two-dimensional real-space image is estimated, and the image is converted into point cloud data based on the depth and color value of each pixel. The three-dimensional virtual space is then generated by a three-dimensional display based on the point cloud data.
[0022] In this way, a virtual recognition object that reproduces an event to be recognized by an image recognition model is constructed based on estimated information ri, making it possible to construct a virtual space that reproduces situations that may occur in real space.
[0023] The virtual space image acquisition unit 14 acquires a virtual space image, which is an image of a virtual space including a virtual recognition target projected from a viewpoint position in the virtual space vs. Specifically, the virtual space image acquisition unit 14 acquires the virtual space image by capturing an image of the virtual recognition target configured in the virtual space with a virtual camera arranged in the virtual space vs.
[0024] 4 is a diagram showing the acquisition of a virtual space image captured by a virtual camera installed in the virtual space and the configuration of training data. As shown in Fig. 4, the virtual space image acquisition unit 14 acquires a virtual space image vp by capturing an image of the virtual space vs in which the virtual recognition target person hm is located using a virtual camera vc installed in the virtual space vs.
[0025] The virtual space image acquisition unit 14 may acquire the virtual space image vp by capturing an image of the virtual space vs reproduced by a virtual space model for representing the real space rs. In this case, a virtual recognition target is placed in the virtual space vs constructed by the model, and the virtual space image can be acquired by capturing an image of the virtual space including the virtual recognition target, so that a virtual space image with higher resolution can be acquired compared to an image captured by optical means.
[0026] Furthermore, the virtual space image acquisition unit 14 may acquire the virtual space image vp by capturing an image of a virtual space in which a virtual recognition target is configured with the real space image rp as the background. That is, by making the placement position of the virtual camera vc in the virtual space vs the placement position of the camera rc in the real space rs when acquiring the real space image rp, the real space image rp can be used as a background other than the virtual recognition target in the virtual space image vp. In this case, the virtual space image acquisition unit 14 may acquire the virtual space image vp by combining the real space image rp as the background with an image of the virtual space in which the recognition target is located. This reduces the processing load required to construct a virtual space model for acquiring the virtual space image.
[0027] It should be noted that the position of the virtual camera vc in the virtual space vs reproduced by the virtual space model does not need to be the same as the position of the camera rc in the real space rs. If the position of the virtual camera vc in the virtual space vs is the same as the position of the camera rc in the real space rs, as described above, the real space image rs can be used as the background in the virtual space image vp, and suitable training data can be obtained for learning an image recognition model that recognizes the image of the real space. On the other hand, even if the position of the virtual camera vc in the virtual space vs is different from the position of the camera rc in the real space rs, the position of the virtual recognition target (e.g., person hp) in the virtual space vs reflects the position of the person hr in the real space rs, and therefore suitable training data can be obtained.
[0028] The training data construction unit 15 outputs training data for the image recognition model, including a virtual space image associated with an event included in the estimated information ri. Specifically, as shown in FIG. 4 , the training data construction unit 15 generates training data td for supervised machine learning by associating data indicating an event to be recognized in image recognition, which is included in the estimated information ri, with the virtual space image vp acquired by the virtual space image acquisition unit 14 as annotation an (correct answer information). In this embodiment, since training data is generated for an image recognition model that recognizes the orientation of a person's face, the annotation an is data such as yaw, pitch, and roll that indicate the orientation of the face.
[0029] The training data composing unit 15 then outputs the generated training data td. Specifically, the training data composing unit 15 may store the generated training data td in the training data storage unit 21.
[0030] In this way, training data td is constructed by annotating events based on the estimated information ri onto a virtual space image vp that projects a virtual space vs including a virtual recognition target reproduced based on the estimated information ri, making it possible to obtain training data with extremely high accuracy as correct information in image recognition.
[0031] The model generation unit 16 generates a second image recognition model through machine learning using training data td. FIG. 5 is a diagram showing the process of generating an image recognition model through machine learning using training data td. As shown in FIG. 5, the model generation unit 16 acquires training data td from the training data storage unit 21, which consists of pairs of virtual space images vp depicting a recognition target that reproduces a predetermined event to be recognized and annotations an consisting of data indicating the predetermined event, and performs machine learning of a second image recognition model md2 using the acquired training data td. The model generation unit 16 generates a second image recognition model md2 that can recognize the predetermined event in the recognition target with high accuracy through machine learning using the training data td.
[0032] FIG. 6 is a flowchart showing the processing contents of an information processing method for generating training data for learning an image recognition model that recognizes events related to a recognition target represented in an image in the information processing device 10, and generating the image recognition model by machine learning using the training data.
[0033] In step S1, the real-space image acquisition unit 11 acquires a real-space image rp obtained by capturing a real space rs including a recognition target (person hr). Specifically, the real-space image acquisition unit 11 acquires a real-space image rp obtained by capturing a real space in which the recognition target exists using a camera rc.
[0034] In step S2, the estimation unit 12 estimates estimated information ri including at least one event based on the real space image rp using a first image recognition model md1 that recognizes an event related to a predetermined recognition target shown in an image.
[0035] In step S3, the virtual space construction unit 13 constructs a virtual recognition target (person hm), which is a virtual recognition target that reproduces a predetermined event, in a virtual space vs that represents real space, based on the estimated information ri.
[0036] In step S4, the virtual space image acquisition unit 14 acquires a virtual space image vp, which is an image obtained by projecting the virtual space vs including the virtual recognition target from the viewpoint position in the virtual space vs.
[0037] In step S5, the training data construction unit 15 constructs and outputs training data td including virtual space images vp associated with the events included in the estimated information ri.
[0038] In step S6, the model generation unit 16 generates a second image recognition model md2 by machine learning using the training data td.
[0039] Next, an information processing program for causing a computer to function as the information processing device 10 of this embodiment will be described with reference to Fig. 7. Fig. 7 is a diagram showing the configuration of the information processing program. The information processing program P1 is configured to include a main module m10 that comprehensively controls information processing in the information processing device 10, a real space image acquisition module m11, an estimation module m12, a virtual space construction module m13, a virtual space image acquisition module m14, a training data construction module m15, and a model generation module m16. Each of the modules m11 to m16 realizes a function for each of the functional units 11 to 16.
[0040] The information processing program P1 may be transmitted via a transmission medium such as a communication line, or may be stored in a recording medium M1 as shown in FIG.
[0041] According to the information processing system 1, information processing device 10, information processing method, and information processing program P1 of the present embodiment described above, a virtual recognition object that reproduces an event recognized by an image recognition model is constructed based on estimated information ri, thereby constructing a virtual space that can occur in real space. Training data td is then constructed from virtual space images vp projecting a virtual space vs that can occur in real space, thereby obtaining training data td suitable for learning an image recognition model to recognize images in real space. Furthermore, since the virtual space images vp included in the training data td are acquired by projecting the virtual space vs, they can be more highly detailed than images optically captured of real space. This makes it possible to improve the accuracy of machine learning for the image recognition model. Furthermore, events based on the estimated information ri are annotated onto virtual space images vp projecting a virtual space vs including a virtual recognition object reproduced based on estimated information ri to construct the training data td, making it possible to obtain training data td with extremely high accuracy as correct answer information in image recognition.
[0042] The information processing device and information processing method according to the present disclosure may have the following configurations: The actions and effects of each configuration will be described as follows.
[0043] An information processing device according to one aspect of the present disclosure includes an estimation unit that estimates estimated information including at least one event based on a real space image using a first image recognition model that recognizes an event related to a recognition target represented in an image, a virtual space construction unit that constructs a virtual recognition target, which is a virtual recognition target that reproduces the event, in a virtual space representing real space based on the estimated information, a virtual space image acquisition unit that acquires a virtual space image, which is an image obtained by projecting the virtual space including the virtual recognition target from a viewpoint position in the virtual space, and a training data construction unit that outputs training data for the image recognition model including a virtual space image associated with the event included in the estimated information.
[0044] An information processing method according to one aspect of the present disclosure includes an estimation step executed by a processor, inferring estimated information including at least one event based on a real-space image captured of a real space including the object of recognition using a first image recognition model that recognizes an event related to the object of recognition represented in an image; a virtual space construction step that constructs a virtual object of recognition in a virtual space, which is a virtual object of recognition that reproduces the event, based on the estimated information; a virtual space image acquisition step that acquires a virtual space image that is an image projected of a virtual space including the virtual object of recognition from a viewpoint position within the virtual space; and a training data construction step that outputs training data for the image recognition model, which includes a real-space image associated with the event included in the estimated information.
[0045] According to the above aspect, a virtual recognition object that reproduces an event recognized by an image recognition model is constructed based on estimated information, thereby constructing a virtual space that could occur in real space. Training data is then constructed from virtual space images projected from the virtual space that could occur in real space, thereby obtaining training data suitable for learning an image recognition model to recognize images of real space. Furthermore, since the virtual space images included in the training data are acquired by projecting the virtual space, they can be more highly detailed than images optically captured from the real space. This makes it possible to improve the accuracy of machine learning for the image recognition model. Furthermore, training data is constructed by annotating events based on the estimated information onto virtual space images projected from a virtual space including a virtual recognition object reproduced based on estimated information, thereby making it possible to obtain training data with extremely high accuracy as correct answer information in image recognition.
[0046] Furthermore, an information processing device according to another aspect may further include a model generation unit that generates a second image recognition model by machine learning using training data.
[0047] According to the above aspects, the accuracy of correct answer information in image recognition is high, and machine learning is performed using training data that is suitable for recognizing images of real space that include the recognition target, making it possible to obtain an image recognition model that can recognize events related to the recognition target with high accuracy.
[0048] In addition, in an information processing device relating to another aspect, the virtual space image acquisition unit may acquire a virtual space image by capturing an image of a virtual recognition target configured in the virtual space using a virtual camera arranged in the virtual space.
[0049] According to the above aspect, since a virtual space image including a virtual recognition target can be acquired by data processing without optically capturing an image of the real space, training data can be easily constructed.
[0050] In addition, in an information processing device according to another aspect, the virtual space image acquisition unit may acquire the virtual space image by capturing an image of a virtual space in which a virtual recognition target is configured with a real space image as a background.
[0051] According to the above aspect, by making the position (viewpoint position) of the virtual camera in the virtual space the same as the position of the camera in capturing the real-space image in the real space, the real-space image can be used as the background of the virtual recognition target in the virtual-space image, thereby reducing the processing load for constructing a virtual-space model for acquiring the virtual-space image.
[0052] In addition, in an information processing device relating to another aspect, the virtual space image acquisition unit may acquire the virtual space image by capturing an image of a virtual space in which a virtual recognition target is placed within a virtual space model that reproduces a real space.
[0053] According to the above aspect, a virtual recognition target is placed in a virtual space constructed as a model, and a high-definition virtual space image can be acquired by capturing an image of the virtual space including the virtual recognition target.
[0054] In an information processing device according to another aspect, the recognition target may be a person, the event may be a facial orientation of the person, and the estimated information may include information indicating the facial orientation of the person.
[0055] According to the above aspect, it is possible to obtain training data for generating an image recognition model that can recognize the direction of a person's face with high accuracy.
[0056] The block diagram shown in FIG. 1 shows functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (e.g., wired, wireless, etc.) and these multiple devices. The functional block may also be realized by combining software with the single device or multiple devices.
[0057] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0058] For example, the information processing device 10 according to an embodiment of the present invention may function as a computer. Fig. 8 is a diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. The information processing device 10 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0059] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the information processing device 10 may be configured to include one or more of the apparatuses shown in FIG. 8, or may be configured to exclude some of the apparatuses.
[0060] Each function of the information processing device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations and control communication via the communication device 1004 and the reading and / or writing of data in the memory 1002 and storage 1003.
[0061] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the functional units 11 to 16 shown in FIG. 1 may be realized by the processor 1001.
[0062] The processor 1001 also reads programs (program codes), software modules, and data from the storage 1003 and / or the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the functional units 11 to 16 of the information processing device 10 may be implemented by a control program stored in the memory 1002 and running on the processor 1001. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0063] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), and a random access memory (RAM). The memory 1002 may also be called a register, a cache, a main memory (primary storage device), or the like. The memory 1002 can store executable programs (program codes), software modules, and the like for implementing an information processing method according to one embodiment of the present invention.
[0064] Storage 1003 is a computer-readable recording medium, and may be, for example, at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including memory 1002 and / or storage 1003.
[0065] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via a wired and / or wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.
[0066] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0067] Furthermore, each device such as the processor 1001 and the memory 1002 is connected to a bus 1007 for communicating information. The bus 1007 may be configured as a single bus, or may be configured as different buses between the devices.
[0068] The information processing device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented by at least one of these pieces of hardware.
[0069] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0070] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (New Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), or other suitable systems, and next-generation systems enhanced based on these. Furthermore, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A with 5G, etc.) may also be applied.
[0071] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0072] In the present disclosure, a specific operation described as being performed by a base station may be performed by its upper node in some cases. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal may be performed by at least one of the base station and another network node other than the base station (for example, an MME or an S-GW, etc., but are not limited to these). Although the above example illustrates a case where there is one other network node other than the base station, a combination of multiple other network nodes (for example, an MME and an S-GW) may also be used.
[0073] Information etc. may be output from a higher layer (or a lower layer) to a lower layer (or a higher layer), or may be input / output via multiple network nodes.
[0074] Input and output information may be stored in a specific location (for example, memory) or managed in a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0075] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0076] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0077] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0078] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0079] Software, instructions, etc. may also be transmitted or received over a transmission medium. For example, if the software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cable, fiber optic cable, twisted pair, and Digital Subscriber Line (DSL), and / or wireless technologies such as infrared, radio, and microwave, these wired and / or wireless technologies are included within the definition of transmission media.
[0080] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0081] It should be noted that terms explained in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meanings.
[0082] As used in this disclosure, the terms "system" and "network" are used interchangeably.
[0083] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed as absolute values, relative values from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.
[0084] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0085] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0086] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly specified otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0087] When designations such as "first," "second," etc. are used in this disclosure, any reference to an element does not generally limit the quantity or order of those elements. These designations may be used herein as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed therein or that the first element must precede the second element in some way.
[0088] The "means" in the configuration of each of the above devices may be replaced with "part," "circuit," "device," etc.
[0089] To the extent that the terms "include," "including," and variations thereof are used herein or in the claims, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, the term "or," as used herein or in the claims, is not intended to be an exclusive or.
[0090] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0091] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0092] The information processing device 10 and information processing method disclosed herein may have the following configuration. [1] An information processing device comprising: an estimation unit that estimates estimated information including at least one event based on the real space image using a first image recognition model that recognizes an event related to a recognition target represented in an image; a virtual space construction unit that constructs a virtual recognition target, which is a virtual recognition target that reproduces the event, in a virtual space representing the real space based on the estimated information; a virtual space image acquisition unit that acquires a virtual space image, which is an image obtained by projecting the virtual space including the virtual recognition target from a viewpoint position in the virtual space; and a training data construction unit that outputs training data for the image recognition model, the virtual space image being associated with the event included in the estimated information. [2] The information processing device described in [1], further comprising: a model generation unit that generates a second image recognition model by machine learning using the training data. [3] The information processing device described in [1] or [2], wherein the virtual space image acquisition unit acquires the virtual space image by capturing an image of the virtual recognition target constructed in the virtual space with a virtual camera disposed in the virtual space. [4] The information processing device according to [3], wherein the virtual space image acquisition unit acquires the virtual space image by capturing an image of a virtual space in which the virtual recognition target is configured with the real space image as a background. [5] The information processing device according to [3], wherein the virtual space image acquisition unit acquires the virtual space image by capturing an image of the virtual space in which the virtual recognition target is placed in a virtual space model that reproduces the real space. [6] The information processing device according to any one of [1] to [5], wherein the recognition target is a person, the event is a facial orientation of the person, and the estimated information includes information indicating a facial orientation of the person.[7] An information processing method executed by a processor, comprising: an estimation step of estimating estimated information including at least one of the events based on a real space image captured of a real space including the recognition target using a first image recognition model that recognizes an event related to the recognition target represented in an image; a virtual space construction step of constructing a virtual recognition target in a virtual space, which is a virtual recognition target that reproduces the event, based on the estimated information; a virtual space image acquisition step of acquiring a virtual space image that is an image projected of the virtual space including the virtual recognition target from a viewpoint position within the virtual space; and a training data construction step of outputting training data for the image recognition model, which includes the real space image associated with the event included in the estimated information.
[0093] 1...information processing system, 10...information processing device, 11...real space image acquisition unit, 12...estimation unit, 13...virtual space construction unit, 14...virtual space image acquisition unit, 15...training data construction unit, 16...model generation unit, 21...training data storage unit, M1...recording medium, m11...real space image acquisition module, m12...estimation module, m13...virtual space construction module, m14...virtual space image acquisition module, m15...training data construction module, m16...model generation module, md1...first image recognition model, md2...second image recognition model, P1...information processing program, rc...camera, ri...estimated information, rp...real space image, td...training data, vc...virtual camera, vs...virtual space.
Claims
1. An information processing device comprising: an estimation unit that estimates estimated information including at least one of the events based on the real space image using a first image recognition model that recognizes an event related to a recognition target represented in an image; a virtual space construction unit that constructs a virtual recognition target, which is a virtual recognition target that reproduces the event, in a virtual space representing the real space based on the estimated information; a virtual space image acquisition unit that acquires a virtual space image, which is an image projected from a viewpoint position within the virtual space and includes the virtual recognition target; and a training data construction unit that outputs training data for the image recognition model, which includes the virtual space image associated with the event included in the estimated information.
2. The information processing device according to claim 1, further comprising a model generation unit that generates a second image recognition model by machine learning using the training data.
3. The information processing device according to claim 1, wherein the virtual space image acquisition unit acquires the virtual space image by capturing an image of the virtual recognition target configured in the virtual space using a virtual camera arranged in the virtual space.
4. The information processing device according to claim 3, wherein the virtual space image acquisition unit acquires the virtual space image by capturing an image of a virtual space in which the virtual recognition target is configured with the real space image as a background.
5. The information processing device according to claim 3, wherein the virtual space image acquisition unit acquires the virtual space image by capturing an image of the virtual space in which the virtual recognition target is placed within a virtual space model that reproduces the real space.
6. The information processing device according to claim 1, wherein the recognition target is a person, the event is the facial orientation of the person, and the estimated information includes information indicating the facial orientation of the person.
7. An information processing method executed by a processor, comprising: an estimation step of estimating estimated information including at least one of the events based on a real space image captured of a real space including the recognition target using a first image recognition model that recognizes an event related to the recognition target represented in an image; a virtual space construction step of constructing a virtual recognition target in a virtual space, which is a virtual recognition target that reproduces the event, based on the estimated information; a virtual space image acquisition step of acquiring a virtual space image that is an image projected of the virtual space including the virtual recognition target from a viewpoint position within the virtual space; and a training data construction step of outputting training data for the image recognition model, which includes the real space image associated with the event included in the estimated information.
Citation Information
Patent Citations
Recognition model distribution system and recognition model updating method
JP2021043622A
Learning data creation method, learning data creation device, and program
JP2022024189A