Image processing apparatus, image processing method, and program

By using semantic segmentation technology to identify real objects and select and control the display area and mode of virtual objects, the problem of lack of personalization in virtual object display in AR image processing is solved, and the user experience is improved.

CN114207671BActive Publication Date: 2025-10-17SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202080055852.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-09
Filing Date
2020-07-07
Publication Date
2025-10-17
Estimated Expiration
2040-07-07

AI Technical Summary

Technical Problem

In existing AR image processing technology, the display processing of virtual objects lacks personalization and fun, resulting in a monotonous user experience.

Method used

Real objects are identified through semantic segmentation technology, and the display area and mode of virtual objects are selected and controlled based on the identification results, realizing personalized superposition display of real objects and virtual objects.

Benefits of technology

It improves the personalization and fun of AR image processing and enhances the user's immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114207671B_ABST
    Figure CN114207671B_ABST
Patent Text Reader

Abstract

Provided are an apparatus and method for selecting a virtual object to be displayed and changing a display mode according to a real object type that is a target area for displaying a virtual object. The apparatus has an object recognition unit that performs processing for recognizing a real object in a real world, and a content display control unit that generates an augmented reality (AR) image in which a real object and a virtual object are superimposed and displayed. The object recognition unit recognizes a real object in a target area that is an area for displaying a virtual object, and the content display control unit performs processing for selecting a virtual object to be displayed and processing for changing a display mode according to an object recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an image processing apparatus, an image processing method, and a program. More specifically, the present disclosure relates to an image processing apparatus, an image processing method, and a program that generate and output an augmented reality (AR) image in which virtual content such as a character image is superimposed and displayed on a real object that can be actually observed. BACKGROUND

[0002] An image in which a virtual object is superimposed and displayed on a real object that can be observed in a real space or a real object and an image is called an augmented reality (AR) image.

[0003] There are various types of virtual objects used in content and games using an AR image, and for example, a virtual object behaving like a human, that is, a character is often used.

[0004] An AR image is displayed by using, for example, a head-mounted display (HMD) worn on a user's eyes, a mobile terminal such as a smart phone, or the like.

[0005] By viewing an AR image, a user can enjoy a feeling as if, for example, a character displayed in an AR image exists in a real world.

[0006] For example, in a case where a character is displayed in an AR image, display processing according to a content reproduction program such as a game program is executed.

[0007] Specifically, in a case where a character output condition recorded in a program is satisfied, a character is displayed in a process defined in the program.

[0008] However, when such a character display defined in a program is executed, similar processing is always repeated, and thus enjoyment is reduced.

[0009] Meanwhile, in recent years, research and use of semantic segmentation as a technique for recognizing objects in an image have been developed. Semantic segmentation is a technique for recognizing types of various objects (for example, a person, a car, a building, a road, a tree, and the like) included in an image captured by an imaging device.

[0010] Note that, for example, Patent Literature 1 (Japanese Patent Application Laid-Open No. 2015-207291) discloses semantic segmentation.

[0011] LIST OF CITATIONS

[0012] PATENT LITERATURE

[0013] Patent Literature 1: Japanese Patent Application Laid-Open No. 2015-207291 SUMMARY

[0014] Problems to be Solved by the Invention

[0015] The present disclosure provides an image processing apparatus, an image processing method, and a program that control a role to be displayed according to a type of a real object.

[0016] Embodiments of the present disclosure provide an image processing apparatus, an image processing method, and a program that recognize a background object of an image through an object recognition process such as the above-described semantic segmentation and control a role to be displayed according to a recognition result.

[0017] Solution to the problem

[0018] A first aspect of the present disclosure is to

[0019] An image processing apparatus comprising:

[0020] an object recognition unit that performs a recognition process of a real object in a real world; and

[0021] a content display control unit that generates an augmented reality (AR) image in which the real object and a virtual object are superimposedly displayed, wherein

[0022] the object recognition unit

[0023] performs an object recognition process that recognizes a real object in a display area of a virtual object, and

[0024] the content display control unit

[0025] selects a virtual object to be displayed according to an object recognition result recognized in the object recognition unit.

[0026] Further, a second aspect of the present disclosure is to

[0027] An image processing method performed in an image processing apparatus, the method comprising:

[0028] an object recognition processing step performed by an object recognition unit, the object recognition processing step performing a recognition process of a real object in a real world;

[0029] a content display control step performed by a content display control unit, the content display control step generating an augmented reality (AR) image in which the real object and a virtual object are superimposedly displayed, wherein

[0030] the object recognition processing step

[0031] is a step that performs an object recognition process that recognizes a real object in a display area of a virtual object, and

[0032] the content display control step

[0033] a step of selecting a virtual object to be displayed according to an object recognition result recognized in the object recognition processing step.

[0034] Further, a third aspect of the present disclosure is to provide

[0035] A program for causing an image processing apparatus to execute image processing, the program:

[0036] causing an object recognition unit to execute an object recognition processing step that performs a recognition process of a real object in a real world;

[0037] causing a content display control unit to execute a content display control step that generates an augmented reality (AR) image in which a real object and a virtual object are superimposed and displayed;

[0038] in the object recognition processing step,

[0039] causing an object recognition process to be executed that recognizes a real object in a display area of a virtual object; and

[0040] in the content display control step,

[0041] causing a step to be executed that selects a virtual object to be displayed according to an object recognition result recognized in the object recognition processing step.

[0042] Note that the program of the present disclosure is, for example, a program that can be provided to a computer system or an information processing apparatus that can execute various program codes by a storage medium or a communication medium provided in a computer-readable form. By providing such a program in a computer-readable form, processing is realized according to the program on the information processing apparatus or the computer system.

[0043] Further other objects, features and merits of the present disclosure will become apparent from the following description of embodiments based on the accompanying drawings. Note that in this specification, the term "system" refers to a logical group configuration of a plurality of apparatuses, and is not limited to a system in which each configured apparatus is in the same housing.

[0044] According to the configuration of the embodiments of the present disclosure, an apparatus and a method in which selection or display mode change of a virtual object to be displayed is performed according to a real object type in a target area to be a display area of the virtual object are realized.

[0045] Specifically, for example, including: an object recognition unit that performs recognition processing of a real object in a real world; and a content display control unit that generates an AR image in which a real object and a virtual object are superimposed and displayed. The object recognition unit recognizes a real object in a target region to be a display region of a virtual object, and the content display control unit performs processing of selecting a virtual object to be displayed or processing of changing a display mode according to an object recognition result.

[0046] With this configuration, a device and a method of performing selection or display mode change of a virtual object to be displayed according to a real object type in a target region to be a display region of a virtual object are realized.

[0047] Note that the advantageous effects described in this specification are merely examples and the advantageous effects of the present technology are not limited to these effects and can include additional effects. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 is a diagram that explains a configuration example of an image processing device of the present disclosure and processing to be performed.

[0049] Figure 2 is a diagram that explains processing performed by the image processing device of the present disclosure.

[0050] Figure 3 is a diagram that explains processing performed by the image processing device of the present disclosure.

[0051] Figure 4 is a diagram that explains processing performed by the image processing device of the present disclosure.

[0052] Figure 5 is a diagram that explains a configuration example of an image processing device of the present disclosure and processing to be performed.

[0053] Figure 6 is a diagram that explains a configuration example of an image processing device of the present disclosure and processing to be performed.

[0054] Figure 7 is a diagram that explains processing performed by the image processing device of the present disclosure.

[0055] Figure 8 is a diagram that explains processing performed by the image processing device of the present disclosure.

[0056] Figure 9 is a diagram that explains a configuration example of an image processing device of the present disclosure.

[0057] Figure 10 is a diagram that explains a data configuration example of a space map data.

[0058] Figure 11FIG. 1 is a diagram illustrating an example of a data configuration of category association update time data.

[0059] Figure 12 FIG. 2 is a diagram illustrating an example of a data configuration of category association virtual object data.

[0060] Figure 13 FIG. 3 is a diagram showing a flowchart illustrating a sequence of processes performed by the image processing apparatus of the present disclosure.

[0061] Figure 14 FIG. 4 is a diagram showing a flowchart illustrating a sequence of target region determination processes performed by the image processing apparatus of the present disclosure.

[0062] Figure 15 FIG. 5 is a diagram showing a flowchart illustrating a sequence of target region determination processes performed by the image processing apparatus of the present disclosure.

[0063] Figure 16 FIG. 6 is a diagram showing a flowchart illustrating a sequence of processes performed by the image processing apparatus of the present disclosure.

[0064] Figure 17 FIG. 7 is a diagram showing a flowchart illustrating a sequence of processes performed by the image processing apparatus of the present disclosure.

[0065] Figure 18 FIG. 8 is a diagram illustrating an example of a hardware configuration of the image processing apparatus of the present disclosure. DETAILED DESCRIPTION

[0066] Hereinafter, the image processing apparatus, the image processing method, and the program of the present disclosure will be described in detail with reference to the accompanying drawings. Note that the description will be given in accordance with the following items.

[0067] 1. Outline of processes performed by the image processing apparatus of the present disclosure

[0068] 2. Example of configuration of the image processing apparatus of the present disclosure

[0069] 3. Sequence of processes performed by the image processing apparatus of the present disclosure

[0070] 3-(1) Basic process sequence performed by the image processing apparatus

[0071] 3-(2) Process sequence of setting a target region in a substantially horizontal surface

[0072] 3-(3) Update sequence of real object recognition process

[0073] 4. Example of hardware configuration of the image processing apparatus

[0074] 5. Summary of the configuration of the present disclosure

[0075] [1. Overview of processing performed by the image processing apparatus of the present disclosure]

[0076] First, referring to Figure 1 and subsequent drawings, an overview of the processing performed by the image processing apparatus of the present disclosure will be described.

[0077] Figure 1 A head-mounted display (HMD) type see-through AR image display device 10 is shown as an example of the image processing apparatus of the present disclosure.

[0078] A user wears the head-mounted display (HMD) type see-through AR image display device 10 to cover the user's eyes.

[0079] The see-through AR image display device 10 includes a see-through display unit (display). The see-through display unit (display) is worn by the user to be disposed at a position in front of the user's eyes.

[0080] The user can observe an external real object as is via the see-through display unit (display) of the see-through AR image display device 10.

[0081] Further, a virtual object image of a virtual object, such as a character image, etc., is displayed on the see-through display unit (display).

[0082] The user can observe the external real object and the virtual object image of the character, etc., together via the see-through AR image display device 10, and can feel as if the virtual object such as the character exists in the real world.

[0083] Figure 1 The right side of FIG. 1 shows an example of an image that the user can observe via the see-through AR image display device 10.

[0084] In (a) observation image example 1, a transmission observation image 21 including an external real object observed via the see-through AR image display device 10 is included. In image example 1, no virtual object is displayed.

[0085] On the other hand, (b) observation image example 2 is an image example in which a virtual object image 22 such as a character image is displayed together with a transmission observation image 21 including an external real object observed via the see-through AR image display device 10. Image example 2 is an image in which the user can observe a real object and a virtual object together.

[0086] As described above, an image in which a virtual object is superimposed and displayed on a real object that can be observed in a real space or a real object and an image is referred to as an augmented reality (AR) image.

[0087] The image processing apparatus of the present disclosure is an apparatus that performs display processing of an AR image.

[0088] The image processing device disclosed herein, for example Figure 1 The illustrated light-transmitting AR image display device 10 performs display control of a virtual object in AR image display.

[0089] The image processing device of the present disclosure, for example Figure 1 The specific processing performed by the illustrated light-transmitting AR image display device 10 is, for example, the following processing.

[0090] (a) A process of generating a three-dimensional map of the real world observed by the user via the translucent AR image display device 10 by applying a simultaneous localization and mapping (SLAM) process that performs self-position estimation and environmental three-dimensional map generation.

[0091] (b) A process of recognizing objects included in the real world through object recognition processing such as semantic segmentation.

[0092] (c) A process of selecting a virtual object (eg, a character) to be displayed and controlling a display mode of the virtual object based on a result of object recognition in the real world.

[0093] For example, the image processing apparatus of the present disclosure performs these types of processing.

[0094] Reference Figure 2 , and subsequent drawings, an example of display control of a virtual object performed by the image processing apparatus of the present disclosure will be described.

[0095] Figure 2 The diagram shown on the left side shows a state in which a user wearing the light-transmitting AR image display device 10 is walking around a park and performs “pointing” at the pond water surface while looking at the pond water surface near a pond.

[0096] The light-transmitting AR image display device 10 includes a camera, captures an image in which the user has performed pointing, and inputs the captured image to a three-dimensional map generation unit via an image analysis unit inside the device.

[0097] The image analysis unit extracts feature points from the captured image, and the three-dimensional map generation unit generates a three-dimensional map of the real world by using the feature points extracted from the image.

[0098] For example, the process of generating a three-dimensional map is performed in real time by a simultaneous localization and mapping (SLAM) process.

[0099] Simultaneous Localization and Mapping (SLAM) is a process that can simultaneously perform self-position estimation and environmental three-dimensional map generation in parallel.

[0100] Furthermore, the three-dimensional map generated by the three-dimensional map generation unit is input to the object recognition unit of the image processing device.

[0101] The object recognition unit uses the three-dimensional map generated by the three-dimensional map generation unit to determine a real object area in the user's pointing direction as the target area 11. In addition, the object in the target area 11 is recognized.

[0102] For example, the object recognition unit performs recognition processing of objects in the real world by applying semantic segmentation processing.

[0103] Semantic segmentation is a type of image recognition processing that uses deep learning to identify objects in an image at the pixel level. Semantic segmentation is a technique for identifying which object category each pixel in an image belongs to based on the degree of matching between dictionary data (learning data) for object recognition, which contains information about the shapes and other characteristics of various real-world objects, and objects in an image captured by a camera, for example.

[0104] Through semantic segmentation, it is possible to identify the types of various objects included in an image captured by a camera, such as a person, a car, a building, a road, a tree, a pond, a lawn, etc.

[0105] exist Figure 2 In the illustrated example, the image analysis unit of the light-transmitting AR image display device 10 determines that the object in the target area 11 in the pointing direction of the user is a “pond” based on the image captured by the camera.

[0106] Furthermore, the content display control unit that performs virtual object display control in the light-transmitting AR image display device 10 inputs the object recognition result of the target area 11 and determines and displays the selection process and display mode of the virtual object (character, etc.) to be displayed according to the object recognition result.

[0107] exist Figure 2 In the example shown, Figure 2 As shown in observation image example 2 (b) on the right, control is performed to display an image of a “water elf character” as the virtual object image 22 .

[0108] This is display control based on the analysis result that the real object in the user's pointing direction is a "pond."

[0109] That is, the content display control unit performs processing to select and display the "water elf character" as the optimal virtual object based on the object recognition result = "pond" of the target area 11.

[0110] Note that the virtual object image 22 is displayed as, for example, a 3D content image.

[0111] Figure 3 Fig. 1 is a diagram illustrating an example of display control of different virtual objects performed by the image processing apparatus of the present disclosure.

[0112] With Figure 2 similarities, Figure 3 The left side of the diagram is a diagram in which a user wearing the see-through AR image display device 10 walks around a park. The user is looking at a lawn in the park while pointing in the direction of the lawn.

[0113] In this case, the object recognition unit of the see-through AR image display device 10 outputs an analysis result that the object in the target area 11, which is the pointing direction of the user, is "lawn" through analysis of the image captured by the imaging device.

[0114] Further, the content display control unit of the see-through AR image display device 10 inputs the object recognition result of the target area 11, and determines and displays the selection process of the virtual object (character, etc.) to be displayed and the display mode in accordance with the object recognition result.

[0115] In Figure 3 the example shown in (b) of the right side of the observation image example 2, control is performed to display an image of a "lawn fairy character" as a virtual object image 23. Figure 3

[0116] This is display control based on the analysis result that the real object in the pointing direction of the user is "lawn".

[0117] That is, the content display control unit performs a process of selecting and displaying a "lawn fairy character" as the best virtual object in accordance with the object recognition result of the target area 11 = "lawn".

[0118] Figure 4 Fig. 2 is a diagram illustrating an example of display control of different virtual objects performed by the image processing apparatus of the present disclosure.

[0119] With Figure 2 and Figure 3 similarities, Figure 4 The left side of the diagram is a diagram in which a user wearing the see-through AR image display device 10 walks around a park. The user is looking at a tree in the park while pointing in the direction of the tree.

[0120] In this case, the object recognition unit of the see-through AR image display device 10 outputs an analysis result that the object in the target area 11, which is the pointing direction of the user, is "tree" through analysis of the image captured by the imaging device.

[0121] ​Further, the content display control unit of the see-through AR image display device 10 inputs the object recognition result of the target region 11, and determines and displays the selection processing and the display mode of the virtual object (character, etc.) to be displayed according to the object recognition result.

[0122] In Figure 4 the example shown in FIG. 1, as Figure 4 the right side (b) observation image example 2, control is executed to display the image of the "tree fairy character" as the virtual object image 23.

[0123] This is display control based on the analysis result that the real object in the pointing direction of the user is "tree".

[0124] That is, the content display control unit executes the processing of selecting and displaying the "tree fairy character" as the best virtual object according to the object recognition result of the target region 11 = "tree".

[0125] As described above, the image processing apparatus of the present disclosure generates a three-dimensional map in the real world by performing three-dimensional shape analysis in the real world using SLAM processing or the like, and in addition, recognizes the object in the target region in the three-dimensional map in the real world by object recognition processing such as semantic segmentation, and performs display control on the virtual object such as a character to be displayed according to the recognition result.

[0126] Note that the target of the real object to be analyzed by the object recognition processing such as semantic segmentation can be processed only in a limited region (i.e., the target region) specified by the user's finger, for example. Therefore, high-speed processing is achieved by limiting the analysis range.

[0127] Note that the image processing apparatus of the present disclosure is not limited to the see-through AR image display device 10 of the head-mounted display (HMD) type described above, and can be configured by an apparatus including various display units. Figure 1 The see-through AR image display device 10 of the head-mounted display (HMD) type described above, and can be configured by an apparatus including various display units.

[0128] For example, the image display type AR image display device 30 shown in FIG. 2 can be used. Figure 5 The image display type AR image display device 30 shown in FIG. 2 includes an image capturing device 31. Figure 5 The image display type AR image display device 30 shown in FIG. 2 includes a non-transmissive display unit (display). The see-through display unit (display) is worn by the user to set a position in front of the user's eyes, which can be observed by the user.

[0129] The image captured by the image capturing device 31 integrated with the image display type AR image display device 30 (i.e., the image 21 shown in FIG. 1) is displayed on the display unit (display) of the image display type AR image display device 30. Figure 5The image 32) captured by the camera is displayed on the display unit in front of the user's eyes. That is, the image of the real object captured by the camera 31 is displayed on the display unit in front of the user's eyes, and the user can confirm the outside scene by watching the image 32 captured by the camera.

[0130] Further, the virtual object image 22 of the virtual object such as a character image and the like is displayed on the display unit (display).

[0131] The user can observe the virtual object image 22 of the character and the like together with the image 32 (i.e., the real object image) captured by the camera displayed on the display unit (display) of the image display type AR image display device 30, and can feel as if the virtual object such as the character exists in the real world.

[0132] Further, the image processing device of the present disclosure can be a portable display device such as Figure 6 The smart phone 40 shown.

[0133] Figure 6 The smart phone 40 shown includes a display unit and a camera 41. The image captured by the camera 41 (i.e., the image 42 captured by the camera shown in the figure) is displayed on the display unit. That is, the image of the real object captured by the camera 41 is displayed on the display unit, and the user can confirm the outside scene by watching the image captured by the camera.

[0134] Further, the virtual object image (e.g., a character image and the like) of the virtual object is displayed on the display unit (display).

[0135] The user can observe the virtual object image 22 of the character and the like together with the image 32 (i.e., the real object image) captured by the camera displayed on the display unit (display) of the image display type AR image display device 30, and can feel as if the virtual object such as the character exists in the real world.

[0136] Note that, in the example of the smart phone 40, in the case where the user touches a specific position of the display unit of the smart phone 40, the image analysis unit of the image processing device (smart phone) analyzes the touch position and further determines the type of the real object at the touch position. Thereafter, the content display control unit of the image processing device (smart phone) performs display control of the virtual object such as a character in accordance with the determination result.

[0137] As described above, the image processing device of the present disclosure performs object recognition, for example, whether the real object in the target region to be a display position of a virtual object is water, grass, or a tree, and performs processing of selecting and displaying a virtual object such as a character to be displayed in accordance with the recognition result.

[0138] Further, the image processing apparatus of the present disclosure not only performs selection processing of a virtual object to be displayed in accordance with the recognition result of a real object in a target region, but also performs processing of changing a display mode of a virtual object such as a character in accordance with the real object recognition result.

[0139] Referring to Figure 7 and subsequent drawings, a specific example of the processing of changing the display mode of a virtual object in accordance with the real object recognition result will be described.

[0140] Figure 7 A display example of an AR image in which a virtual object (character) jumps out of a pond as a real object is shown as an example of an AR image displayed on the display unit of the image processing apparatus of the present disclosure.

[0141] The content display control unit of the image processing apparatus of the present disclosure performs processing of moving or making a virtual object such as a character displayed in a target region in accordance with a preset program.

[0142] Figure 7 The example shown shows an example of performing display in which the character displayed in the target region 11 is moved in the upward direction.

[0143] Figure 7 (1), the timing-1 when the virtual object (character) jumps out of the pond is a state in which the upper half of the virtual object (character) is displayed on the water and the lower half is in the water.

[0144] In this state, as shown in Figure 7 (1), the content display control unit of the image processing apparatus displays the virtual object image 50 on the water and the virtual object image 51 in the water in different display modes.

[0145] That is, the virtual object image 50 on the water is displayed as a normal image with a clear outline, and the virtual object image 51 in the water is displayed as an image with three-dimensional distortion because it exists in the water.

[0146] Further, the content sound control unit of the image processing apparatus outputs the sound of the water (splashing sound, etc.) via the speaker as a sound effect when the character moves onto the water.

[0147] Figure 7 (2), the timing-2 when the virtual object (character) jumps out of the pond is a state in which the entire body of the virtual object (character) is displayed on the water. Figure 7 (1) after the AR image display example. At this time, the entire body of the virtual object (character) is displayed above the water. In this state, the virtual object image 50 above the water is an image of the entire character, and the content display control unit of the image processing apparatus displays the entire character image as an image with a clear outline.

[0148] Figure 8 is another specific example illustrating a process of changing a display mode of a virtual object such as a character in accordance with a real object recognition result.

[0149] Figure 8 A display example of an AR image in which a shadow of a virtual object (a character) is displayed is shown as an example of an AR image displayed on the display unit of the image processing apparatus of the present disclosure.

[0150] That is, when the content display control unit displays the character in the target region 11, the shadow of the character is also displayed.

[0151] As Figure 8 indicated in "(1) Display example of a shadow of a character in a case where a surface on which a shadow appears is a flat surface"

[0152] is a display example of a shadow in a case where a surface on which a virtual object (a character) image 50 appears is a flat surface (for example, a floor in a room or a sidewalk outside).

[0153] As described above, in a case where a surface on which a virtual object (a character) image 50 appears is a flat surface, when a virtual object image 50 as a three-dimensional character is displayed in a target region 11, a content display control unit of the image processing apparatus displays a virtual object shadow image 52 indicating a shadow of the virtual object image 50 as an image having a clear outline.

[0154] On the other hand, Figure 8 "(2) Display example of a shadow of a character in a case where a surface on which a shadow appears is a non-flat surface (a sand pit or the like)"

[0155] is a display example of a shadow in a case where a surface on which a virtual object (a character) image 50 appears is not a flat surface (for example, a sand pit).

[0156] As described above, in a case where a surface on which a virtual object (a character) image 50 appears is a non-flat surface such as a sand pit, when a virtual object image 50 as a three-dimensional character is displayed in a target region 11, a content display control unit of the image processing apparatus displays a virtual object shadow image 52 indicating a shadow of the virtual object image 50 as an image having a non-clear outline and being non-flat.

[0157] As described above, a content display control unit of the image processing apparatus of the present disclosure performs control in accordance with a recognition result of a real object in a target region in which a virtual object is displayed to change a display mode of the virtual object, and performs display, and in addition, a content sound control unit performs output control of sound effects in accordance with a recognition result of a real object in a target region.

[0158] [2. Configuration example of the image processing apparatus of the present disclosure]

[0159] Next, a configuration example of the image processing apparatus of the present disclosure will be described.

[0160] As described above, the image processing apparatus of the present disclosure can be implemented as an apparatus having various forms, such as the see-through type AR image display device 10 described with reference to Figure 1 the image display type AR image display device 30 described with reference to Figure 5 the smartphone 40 described with reference to Figure 6 a portable display apparatus.

[0161] Figure 9 is a block diagram showing a configuration example of the image processing apparatus of the present disclosure that can adopt these various forms.

[0162] The configuration of the image processing apparatus 100 shown in Figure 9 will be described.

[0163] As shown in Figure 9 , the image processing apparatus 100 includes a data input unit 110, a data processing unit 120, a data output unit 130, and a communication unit 140.

[0164] The data input unit 110 includes an external imaging camera 111, an internal imaging camera 112, a motion sensor (gyro, acceleration sensor, etc.) 113, an operation unit 114, and a microphone 115.

[0165] The data processing unit 120 includes an external captured image analysis unit 121, a three-dimensional map generation unit 122, an internal captured image analysis unit 123, a device posture analysis unit 124, a sound analysis unit 125, an object recognition unit 126, spatial map data 127, and category association update time data 128.

[0166] The data output unit 130 includes a content display control unit 131, a content sound control unit 132, a display unit 133, a speaker 134, and category association virtual object data (3D model, sound data, etc.) 135.

[0167] The external imaging camera 111 of the data input unit 110 captures an external image. For example, an image of an external scene or the like in an environment in which a user wearing an HMD is present is captured. In the case of a mobile terminal such as a smartphone, a camera included in the smartphone or the like is used.

[0168] The internal imaging camera 112 is basically a component unique to the HMD, and captures an image of a user's eye region to analyze the user's line-of-sight direction.

[0169] The motion sensor (gyro, acceleration sensor, etc.) 113 detects the posture and movement of the image processing apparatus 100 main body (e.g., HMD, smart phone, etc.).

[0170] The motion sensor 113 includes, for example, a gyro, an acceleration sensor, an orientation sensor, a single positioning sensor, an inertial measurement unit (IMU), etc.

[0171] The operation unit 114 is an operation unit that a user can operate, and is used for, for example, input of a target region, input of other processing instructions, etc.

[0172] The microphone 115 is used for input of instructions, etc. through a user's voice input. In addition, the microphone can also be used for input of external environment sound.

[0173] Next, the components of the data processing unit 120 will be described.

[0174] The external captured image analysis unit 121 inputs the external captured image captured by the external imaging camera 111, and extracts a feature point from the external captured image.

[0175] The processing of extracting a feature point is processing of extracting a feature point for generating a three-dimensional map, and the extracted feature point information is input to the three-dimensional map generation unit 122 together with the external captured image captured by the external imaging camera 111.

[0176] The three-dimensional map generation unit 122 generates a three-dimensional map including an external real object based on the external captured image captured by the external imaging camera 111 and the feature point extracted by the external captured image analysis unit 121.

[0177] For example, the processing of generating a three-dimensional map is executed by simultaneous localization and mapping (SLAM) processing as real-time processing.

[0178] As described above, the simultaneous localization and mapping (SLAM) processing is processing capable of performing self position estimation and environment three-dimensional map generation in parallel.

[0179] The three-dimensional map data of the external environment generated by the three-dimensional map generation unit 122 is input to the object recognition unit 126.

[0180] The internal captured image analysis unit 123 analyzes the user's line-of-sight direction based on the image of the user's eye region captured by the internal imaging camera 112. Similar to the internal imaging camera 112 described above, the internal captured image analysis unit 123 is basically a component unique to the HMD.

[0181] The user gaze information analyzed by the internal-captured-image analysis unit 123 is input to the object recognition unit 126.

[0182] The device posture analysis unit 124 analyzes the posture and movement of the image processing apparatus 100 main body such as an HMD or a smart phone, based on sensor detection information measured by a motion sensor (gyro, acceleration sensor, etc.) 113.

[0183] The posture and movement information of the image processing apparatus 100 main body analyzed by the device posture analysis unit 124 is input to the object recognition unit 126.

[0184] The sound analysis unit 125 analyzes the user voice and environmental sound input from the microphone 115. The analysis result is input to the object recognition unit 126.

[0185] The object recognition unit 126 inputs the three-dimensional map generated by the three-dimensional map generation unit 122, determines a target region to be set as a display region of a virtual object, and further executes recognition processing of a real object in the determined target region. The object recognition processing is executed, for example, the target region is a pond, a tree, etc.

[0186] The processing of recognizing the target region can be executed by various methods.

[0187] For example, the above processing can be executed by using an image of a user's finger included in the three-dimensional map.

[0188] An intersection between an extension line in the user pointing direction and a real object on the three-dimensional map is obtained, and for example, a circular region of a predetermined radius centered on the intersection is determined as the target region.

[0189] Note that the designation of the target region can also be executed by a method other than the user's pointing. The object recognition unit 126 can use any of the following information as information for determining the target region.

[0190] (a) The user gaze information analyzed by the internal-captured-image analysis unit 123

[0191] (b) The posture and movement information of the image processing apparatus 100 main body analyzed by the device posture analysis unit 124

[0192] (c) The user operation information input via the operation unit 114

[0193] (d) The user voice information analyzed by the sound analysis unit 125

[0194] In a case where "(a) the user's line-of-sight information analyzed by the internal captured image analysis unit 123" is used, the object recognition unit 126 obtains an intersection between an extension line of the user's line-of-sight direction and a real object on the three-dimensional map and determines, for example, a circular region having a predetermined radius centered on the intersection as the target region.

[0195] In a case where "(b) the posture and movement information of the image processing apparatus 100 main body analyzed by the device posture analysis unit 124" is used, the object recognition unit 126 obtains an intersection between an extension line in a front direction of the HMD worn by the user or the smart phone held by the user and a real object on the three-dimensional map and determines, for example, a circular region having a predetermined radius centered on the intersection as the target region.

[0196] In a case where "(c) the user operation information input via the operation unit 114" is used, the object recognition unit 126 determines the target region based on, for example, the user operation information input via the input unit of the image processing apparatus 100.

[0197] For example, in a case where the above-described Figure 6 In the configuration of the smart phone illustrated in FIG. 10, the following processing can be performed: the user's finger inputs the screen position specification information as the user operation information, and the specified position of the screen position specification information is set as the center position of the target region.

[0198] Note that, in addition to this, a bar-shaped pointing member separate from the image processing apparatus 100 can be used as the operation unit 114, the pointing direction information of the pointing member can be input to the object recognition unit 126, and the target region can be determined based on the pointing direction.

[0199] In a case where "(d) the user's voice information analyzed by the sound analysis unit 125" is used, the object recognition unit 126 analyzes, for example, the user's utterance to determine the target region.

[0200] For example, in a case where the user's utterance is an utterance such as "the pond in front", the pond is determined as the target region.

[0201] Further, the object recognition unit 126 can perform a target region determination processing other than these. For example, a horizontal surface detection processing such as a ground surface, a floor surface, or a water surface can be performed based on the three-dimensional map generated from the image captured by the external imaging camera 111 or the detection information from the motion sensor 113, and a processing of determining a region closest to a center region of the captured image among the horizontal surfaces as the target region can be performed.

[0202] In addition, the following processing can also be performed: for example, the user performs an operation of throwing a virtual ball, the image of the virtual ball is captured by the external imaging camera 111, the landing point of the ball is analyzed by analyzing the captured image, and the landing point is set as the center position of the target area.

[0203] The object recognition unit 126 determines the target area to be used as the virtual object display area by using any of the above methods. In addition, the object recognition processing is performed on the real object in the determined target area. The object recognition processing is performed, for example, the target area is a pond, a tree, etc.

[0204] As described above, for example, the object recognition process for a real object is performed by applying the semantic segmentation process.

[0205] Semantic segmentation is a technology that identifies to which object category each constituent pixel of an image belongs based on the degree of matching between dictionary data (learning data) for object recognition, in which shape information and other feature information of various actual objects are registered, and objects in an image captured by, for example, a camera device.

[0206] Through semantic segmentation, it is possible to identify the types of various objects included in an image captured by a camera, such as a person, a car, a building, a road, a tree, a pond, a lawn, etc.

[0207] Note that the object recognition processing performed by the object recognition unit 126 is performed only for the target area or only for a limited range of the surrounding area including the target area. By performing this processing within a limited range, high-speed processing, i.e., real-time processing, can be performed. Note that, for example, real-time processing means that the recognition processing of the real object is performed immediately after the user specifies the target area. As a result, for example, object recognition is completed without time delay while the user is observing the target area.

[0208] The recognition result of the real object in the target area analyzed by the object recognition unit 126 is input to the content display control unit 131 and the content sound control unit 132 of the data output unit 130 .

[0209] The content display control unit 131 of the data output unit 130 inputs the object recognition result for the target area from the object recognition unit 126, determines the selection processing and display mode of the virtual object (character, etc.) to be displayed according to the object recognition result, and displays the virtual object on the display unit 133.

[0210] Specifically, for example, the previously described Figure 2 to Figure 4 、 Figure 7 and Figure 8 Display processing of the virtual objects (characters, etc.) shown.

[0211] The content sound control unit 132 of the data output unit 130 inputs the object recognition result of the target region from the object recognition unit 126, determines the sound to be output according to the object recognition result, and outputs the sound via the speaker 134.

[0212] Specifically, for example, as described previously Figure 7 As illustrated, in a case where the virtual object appears from the pond that is a real object, the process of outputting the sound of water is executed.

[0213] Note that the content display control unit 131 and the content sound control unit 132 of the data output unit 130 acquire the 3D content of the virtual object and various sound data recorded in the category-associated virtual object data 135 and execute data output.

[0214] In the category-associated virtual object data 135, the 3D content of the virtual object and various sound data for display associated with the real object type (category) corresponding to the recognition result of the real object in the target region are recorded.

[0215] A specific example of the category-associated virtual object data 135 will be described later.

[0216] Further, in a case where the image processing apparatus 100 is configured to execute a process of displaying the image captured by the imaging device, for example, as in the AR image display device 30 that displays the image captured by the imaging device described with reference to Figure 5 or the smart phone 40 described with reference to Figure 6 The content display control unit 131 inputs the captured image of the external imaging device 111, generates a display image in which the virtual object is superimposed on the captured image, and displays the display image on the display unit 133, as in the smart phone 40 described with reference to

[0217] The communication unit 140 communicates with, for example, an external server, and acquires the 3D content of the character that is a virtual content. Further, various data and parameters required for data processing can be acquired from the external server.

[0218] Note that the object recognition unit 126 stores the recognition result as the spatial map data 127 in the storage unit when executing the recognition process of the real object in the target region.

[0219] Figure 10 A data configuration example of the spatial map data 127 is illustrated.

[0220] As Figure 10 illustrated, the association data of each piece of data below is stored in the spatial map data.

[0221] (a) Time stamp (second)

[0222] (b) Position information

[0223] (c) category

[0224] (d) elapsed time (sec) after recognition processing

[0225] (a) timestamp (sec) is time information on execution of object recognition processing.

[0226] (b) position information is position information of a real object that is an object recognition target. As a method of recording position information, various methods can be used. An example shown in the figure is an example described as a grid by a list of three-dimensional coordinates (x, y, z). In addition, for example, position information of a center position of a target region can be recorded.

[0227] (c) category is object type information that is an object recognition result.

[0228] (d) elapsed time (sec) after recognition processing is time elapsed from completion of object recognition processing.

[0229] Note that, after determination of a target region, the object recognition unit 126 immediately executes recognition processing of a real object in the target region, and thereafter, repeatedly executes object recognition processing on the region, and sequentially updates Figure 10 the spatial map data shown.

[0230] However, an interval of the update processing varies according to a type (category) of a recognized real object.

[0231] Designation data of an update time that differs according to a type (category) of a real object is registered as category-associated update time data 128 in advance.

[0232] Figure 11 A data example of the category-associated update time data 128 is shown.

[0233] As Figure 11 shown, the category-associated update time data 128 is data in which the following data are associated with each other.

[0234] (a) ID

[0235] (b) classification

[0236] (c) category

[0237] (d) update time (sec)

[0238] (a) ID is an identifier of registered data.

[0239] (b) classification is a classification of a type (category) of a real object.

[0240] (c) A category is the type information of a real object.

[0241] (d) Update Time (Seconds) is a time indicating an update interval of the real object recognition process.

[0242] For example, in the case of the category (object type) of ID001 = lawn, the update time is 3600 seconds (= 1 hour). In an object such as lawn, as the change in elapsed time is small, the update time is set to be long.

[0243] On the other hand, for example, in the case of category (object type) = shadow with ID = 004, the update time is 2 seconds. In an object such as a shadow, the change with the passage of time is large, so the update time is set short.

[0244] The object recognition unit 126 refers to the data of the category association update time data 128 and repeatedly performs the object recognition process at time intervals defined for the recognized object as needed. Real objects detected by new recognition processes are sequentially registered as reference objects. Figure 10 Described spatial map data 127.

[0245] Furthermore, as described above, the content display control unit 131 and the content sound control unit 132 of the data output unit 130 acquire 3D content and various sound data of virtual objects recorded in the category-associated virtual object data 135 and perform data output.

[0246] In the category-associated virtual object data 135, 3D contents and various sound data of virtual objects for display associated with the real object type (category) corresponding to the recognition result of the real object in the target area are recorded.

[0247] Will refer to Figure 12 A specific example of the category-associated virtual object data 135 will be described.

[0248] like Figure 12 As shown, the following data are recorded in association with each other in the category-associated virtual object data 135 .

[0249] (a) Category

[0250] (b) Virtual object 3D model (character 3D model)

[0251] (c) Output sound

[0252] (a) A category is the type information of a real object.

[0253] As (b) virtual object 3D model (character 3D model), 3D models of virtual objects (characters) to be output (displayed) according to each category, that is, the type of real object in the target region, are registered. Note that in the example shown in the figure, the ID of the 3D model is recorded together with the 3D model, but for example, only the ID can be recorded, and the 3D model associated with the ID can be acquired from another database based on the ID.

[0254] As (c) output sound, sound data to be output according to each category, that is, the type of real object in the target region, is registered.

[0255] As described above, in the category-associated virtual object data 135, 3D contents of virtual objects for display associated with the type (category) of real object corresponding to the recognition result of the real object in the target region and various sound data are recorded.

[0256] Note that the output mode information of each virtual object is also recorded in the category-associated virtual object data 135. For example, as described earlier with reference to Figure 7 and Figure 8 For example, display mode information in the case where the real object in the target region is water, or display mode information in the case where the real object in the target region is a sandpit is also recorded.

[0257] The content display control unit 131 and the content sound control unit 132 of the data output unit 130 acquire the 3D contents of virtual objects and various sound data recorded in the category-associated virtual object data 135 that stores data as shown in Figure 12 and execute data output.

[0258] [3. Sequence of processing performed by the image processing apparatus of the present disclosure]

[0259] Next, the sequence of processing performed by the image processing apparatus 100 of the present disclosure will be described.

[0260] Note that the following described multiple processing sequences will be described in order.

[0261] (1) Basic processing sequence performed by the image processing apparatus

[0262] (2) Sequence of processing of setting a target region in a set region provided in a substantially horizontal surface

[0263] (3) Update sequence of real object recognition processing

[0264] (3-(1) Basic processing sequence performed by the image processing apparatus)

[0265] First, refer to Figure 13Referring to the flowchart shown in FIG. 1 , a sequence of basic processing executed by the image processing apparatus 100 of the present disclosure will be described.

[0266] Note that according to Figure 13 The processes of the flowcharts shown in the following figures are processes mainly executed in the data processing unit 120 of the image processing apparatus 100. The data processing unit 120 includes a CPU having a program execution function, and executes processes according to a flow in accordance with a program stored in a storage unit.

[0267] Next, we will describe Figure 13 Processing of each step of the process shown.

[0268] (Step S101)

[0269] First, in step S101 , the data processing unit 120 of the image processing apparatus 100 inputs a captured image of an external imaging camera.

[0270] (Step S102)

[0271] Next, in step S102 , the data processing unit 120 extracts feature points from a captured image input by an external imaging camera.

[0272] The processing is done by Figure 9 The processing is performed by the external captured image analysis unit 121 of the data processing unit 120 shown.

[0273] The external captured image analysis unit 121 extracts feature points from a captured image input by an external imaging camera. This feature point extraction process is a process for extracting feature points for generating a three-dimensional map. The extracted feature point information is input to the three-dimensional map generation unit 122 along with the external captured image captured by the external imaging camera.

[0274] (Step S103)

[0275] Next, in step S103 , the data processing unit generates a three-dimensional map by using the captured image of the outside captured by the external imaging camera and its feature point information.

[0276] The processing is done by Figure 9 The processing is performed by the three-dimensional map generation unit 122 of the data processing unit 120 shown.

[0277] The three-dimensional map generation unit 122 generates a three-dimensional map including external real objects based on the captured image of the outside captured by the external imaging camera 111 and the feature points extracted by the external captured image analysis unit 121 .

[0278] For example, the process of generating a three-dimensional map is executed as real-time processing by simultaneous localization and mapping (SLAM) processing.

[0279] (Step S104)

[0280] Next, in step S104, the data processing unit executes target region determination processing.

[0281] This processing is processing executed by the object recognition unit 126 of the data processing unit 120 shown in Figure 9

[0282] The object recognition unit 126 determines a target region to be a virtual object display region.

[0283] As described above, various methods can be applied to the target region determination processing.

[0284] For example, the above-described processing can be executed by using an image of a user's finger included in the three-dimensional map.

[0285] That is, an intersection between an extension line in a pointing direction of the user and a real object on the three-dimensional map is obtained, and for example, a circular region having a predetermined radius centered on the intersection is determined as the target region.

[0286] Furthermore, the target region can also be determined by using input information from each component of the data input unit 110 shown in Figure 9

[0287] (a) user line-of-sight information analyzed by the internal captured image analysis unit 123

[0288] (b) posture and movement information of the image processing apparatus 100 main body analyzed by the device posture analysis unit 124

[0289] (c) user operation information input via the operation unit 114

[0290] (d) user voice information analyzed by the sound analysis unit 125

[0291] For example, the target region can be determined by using any of these input information.

[0292] A representative target region determination sequence will be described with reference to Figure 14 and Figure 15

[0293] Figure 14 (1) A target region determination sequence based on analysis of a pointing direction of a user is shown. The target region determination processing based on analysis of a pointing direction of a user is executed in the following processing sequence. ​​​

[0294] First, in step S211, the pointing direction of the user is analyzed. This analysis processing is performed by using the three-dimensional map generated by the three-dimensional map generation unit 122.

[0295] Next, in step S212, the intersection between the straight line formed by the extension of the pointing direction of the user and the real object is detected. This processing is also performed by using the three-dimensional map generated by the three-dimensional map generation unit 122.

[0296] Finally, in step S213, a circular region centered on the intersection between the straight line formed by the extension of the pointing direction of the user and the real object is determined as the target region.

[0297] Note that the shape of the target region is arbitrary, and can be rectangular in addition to circular. The size of the target region is also arbitrary, and can be set to various sizes.

[0298] However, it is preferable to define the shape and size in advance, and determine the target region in accordance with the definition.

[0299] Figure 14 (2) Target region determination sequence based on analysis of user line-of-sight direction. The target region determination processing based on analysis of the user line-of-sight direction is performed in the following processing sequence.

[0300] First, in step S221, the line-of-sight direction of the user is analyzed. This analysis processing is performed by the internal captured image analysis unit 123 based on the captured image by the internal imaging camera 112.

[0301] Next, in step S222, the intersection between the straight line formed by the extension of the line-of-sight direction of the user and the real object is detected. This processing is performed by using the three-dimensional map generated by the three-dimensional map generation unit 122.

[0302] Finally, in step S223, a circular region centered on the intersection between the straight line formed by the extension of the line-of-sight direction of the user and the real object is determined as the target region.

[0303] Note that, as described above, the shape and size of the target region can be set differently.

[0304] Figure 15 (3) Target region determination sequence based on analysis of user operation information. The target region determination processing based on analysis of the user operation information is performed in the following processing sequence.

[0305] First, in step S231, the user operation information is analyzed. For example, the user operation is a touch operation on the smart phone described above with reference to Figure 6 Fig. 6. The user operation information is analyzed by the internal captured image analysis unit 123 based on the captured image by the internal imaging camera 112.

[0306] Next, in step S232, a real object designation position based on the user operation information is detected. This processing is performed, for example, as detection processing of a position at which the user's finger contacts.

[0307] Finally, in step S233, a circular region centered on the real object designation position based on the user operation information is determined as the target region.

[0308] Note that, as described above, the shape and size of the target region can be set differently.

[0309] Figure 15 (4) A target region determination sequence based on analysis of user voice information is shown. The target region determination processing based on analysis of user voice information is performed in the following processing sequence.

[0310] First, in step S241, the user's uttered voice is analyzed. For example, the user's uttered voice such as "the pond in front" is analyzed.

[0311] Next, in step S242, a real object designation position based on the user voice information is detected.

[0312] Finally, in step S243, a circular region centered on the real object designation position based on the user's uttered voice is determined as the target region.

[0313] Note that, as described above, the shape and size of the target region can be set differently.

[0314] In addition to the description of Figure 14 and Figure 15 , for example, the following target region determination processing can be performed.

[0315] (a) Target region determination processing using the posture and movement information of the image processing apparatus 100 main body analyzed by the device posture analysis unit 124.

[0316] (b) Detection processing of a horizontal surface such as a ground, floor surface, or water surface based on a three-dimensional map generated from an image captured by the external imaging camera 111 or detection information from the motion sensor 113, and processing of determining as the target region a region in the horizontal surface that is closest to the center region of the captured image.

[0317] (c) Processing in which the user performs an operation of throwing a virtual ball, an image of the virtual ball is captured by the external imaging camera 111, the landing point of the ball is analyzed by analyzing the captured image, and the landing point is determined as the center position of the target region.

[0318] (d) Additionally, at least any one of a user action, a user line of sight, a user operation, a user position, and a user posture is analyzed, and a target region is determined based on an analysis result.

[0319] Return Figure 13 The flowchart illustrated above will be described with continuation.

[0320] As described above, the object recognition unit 126 of the data processing unit 120 of the image processing apparatus 100 executes the target region determination processing in step S104.

[0321] (Step S105)

[0322] Next, in step S105, the data processing unit recognizes a real object in the target region.

[0323] Specifically, the object recognition processing is executed, for example, the target region is a pond, a tree, or the like.

[0324] As described above, the object recognition processing of the real object is executed, for example, by applying the semantic segmentation processing.

[0325] The semantic segmentation is a technique of identifying to which object classification each constituent pixel of an image belongs based on, for example, a matching degree between dictionary data (learning data) for object recognition in which shape information and other characteristic information of various actual objects are registered and an object in an image captured by, for example, an imaging device.

[0326] Note that the recognition processing of the real object executed by the object recognition unit 126 is executed only for the target region or only for a limited range including a surrounding region of the target region. By performing such processing in a limited range, high-speed processing, that is, real-time processing can be executed.

[0327] The recognition result of the real object in the target region analyzed by the object recognition unit 126 is input to the content display control unit 131 and the content sound control unit 132 of the data output unit 130.

[0328] (Step S106)

[0329] Next, in step S106, based on the real object (category) in the recognized target region, the type and the output mode of the virtual object to be displayed in the target region are determined.

[0330] This processing is processing executed by the content display control unit 131 and the content sound control unit 132 of the data output unit 130 of the image processing apparatus 100 illustrated in Figure 9

[0331] ​The content display control unit 131 and the content sound control unit 132 of the data output unit 130 refer to the category association virtual object data 135 in which the data described above is recorded, to determine the type of virtual object and the output mode to be displayed in the target region. Figure 12

[0332] That is, the following processing and the like is executed: an entry in which the type (category) of real object in the target region is recorded is selected from among the entries of the category association virtual object data 135, and a virtual object recorded in the entry is determined as an output object.

[0333] (Step S107)

[0334] Finally, in step S107, a virtual object is output (displayed) to the target region in accordance with the type of virtual object and the output mode to be displayed in the target region determined in step S106.

[0335] This processing is also processing executed by the content display control unit 131 and the content sound control unit 132 of the data output unit 130 of the image processing apparatus 100 shown in Figure 9

[0336] The content display control unit 131 inputs the object recognition result for the target region from the object recognition unit 126, determines the selection processing of virtual object (character and the like) to be displayed and the display mode in accordance with the object recognition result, and displays the virtual object on the display unit 133.

[0337] Specifically, for example, the display processing of the virtual object (character and the like) shown in Figure 2 to Figure 4 , Figure 7 and Figure 8 is executed as described previously.

[0338] Further, the content sound control unit 132 inputs the object recognition result for the target region from the object recognition unit 126, determines the sound to be output in accordance with the object recognition result, and outputs the sound via the speaker 134.

[0339] Specifically, for example, as shown in Figure 7 described previously, in a case in which a virtual object appears from a pond which is a real object, processing of outputting a water sound is executed.

[0340] (3-(2) Processing sequence of the setting region in which the target region is set in a substantially horizontal surface)

[0341] Next, the sequence of the processing of the setting region in which the target region is set in a substantially horizontal surface will be described with reference to the flowchart shown in Figure 16

[0342] ​​​In a case where a virtual object such as a character is displayed on a real object in the real world, if the virtual object is displayed on the ground in a case where the real object is outdoors, if the virtual object is displayed on the floor in a case where the real object is indoors, a more natural character display becomes possible, and it is possible to make the user feel that the character actually exists in the real world.

[0343] For this purpose, control is executed to set a target region to be an output region of a virtual object (i.e., a character) in a substantially horizontal surface such as the ground or a floor is effective.

[0344] Figure 16 The flowchart illustrated is a flowchart that illustrates a processing sequence of the image processing apparatus 100 that executes such processing.

[0345] Hereinafter, the processing of each step of the flowchart illustrated will be described. Figure 16

[0346] Note that, Figure 16 The processing of steps S101 to S103 and steps S105 to S107 of the flowchart illustrated is processing similar to the processing of the corresponding steps of the basic processing flow described earlier with reference to Figure 13

[0347] Figure 16 The processing in steps S301 to S303 of the flow and the processing in step S104 are points different from the processing described earlier with reference to Figure 13

[0348] The processing of each step will be described.

[0349] (Step S301)

[0350] Step S301 is processing of inputting sensor detection information from the motion sensor 113 of the data input unit 110 of the image processing apparatus 100 illustrated to the device posture analysis unit 124 of the data processing unit 120. Figure 9 As described earlier with reference to

[0351] , the motion sensor 113 includes a gyroscope, an acceleration sensor, and the like, and is a sensor that detects the posture and movement of the image processing apparatus 100 main body, for example, an HMD or a smartphone. Figure 9

[0352] The sensor detection information is input from the motion sensor 113 to the device posture analysis unit 124 of the data processing unit 120.

[0353] (Step S302)

[0354] ​​​​Next, in step S302 , the direction of gravity is estimated based on the motion sensor detection information.

[0355] The processing is done by Figure 9 The processing is performed by the device posture analysis unit 124 of the data processing unit 120 shown.

[0356] The device posture analysis unit 124 of the data processing unit 120 calculates the gravity direction by using sensor detection information from a gyroscope, an acceleration sensor, and the like constituting the motion sensor 113 .

[0357] (Step S303)

[0358] Next, in step S303 , a detection process of a horizontal surface area is performed.

[0359] The processing is done by Figure 9 The processing performed by the object recognition unit 126 is shown.

[0360] The object recognition unit 126 detects a horizontal surface area in the three-dimensional map by using the three-dimensional map generated by the three-dimensional map generation unit 122 and the gravity direction information input from the device posture analysis unit 124. Specifically, for example, the ground, the floor surface, etc. are detected.

[0361] Note that the horizontal surface area to be detected is not limited to a completely horizontal surface, but only needs to be a substantially horizontal area.

[0362] For example, a certain degree of unevenness, a slope with a certain degree of inclination, etc. are also determined and detected as a horizontal surface area.

[0363] What degree of unevenness or inclination is allowed as a horizontal surface area can be preset.

[0364] (Step S104)

[0365] Next, in step S104 , the data processing unit performs target region determination processing.

[0366] However, in the present processing example, the target area is selected only from among the horizontal surface areas detected in step S303 .

[0367] The processing is done by Figure 9 The processing is performed by the object recognition unit 126 of the data processing unit 120 shown.

[0368] The object recognition unit 126 determines a target area to be a virtual object display area only within the horizontal surface area detected in step S303 .

[0369] As described above, various methods can be applied to the target region determination process.

[0370] For example, the above-described processing can be performed by using an image of a user's finger included in the three-dimensional map.

[0371] That is, an intersection between an extension line of a user pointing direction and a horizontal surface region that is a horizontal surface on the three-dimensional map and is determined to be a horizontal surface such as a ground surface or a floor surface is obtained, and a circular region having a predetermined radius centered on the intersection with the horizontal surface is determined as a target region.

[0372] Note that, as data for determining a target region, various types of information as described previously with reference to Figure 13 the flowchart shown in FIG. 6 can be used. For example, the following input information can be used.

[0373] (a) user line-of-sight information analyzed by the internal image analysis unit 123

[0374] (b) posture and movement information of the image processing apparatus 100 main body analyzed by the device posture analysis unit 124

[0375] (c) user operation information input via the operation unit 114

[0376] (d) user voice information analyzed by the sound analysis unit 125

[0377] For example, any of these input information can be used to determine a target region.

[0378] The processing in steps S101 to S103 and the processing in step S105 and subsequent steps are similar to the processing in the flowchart shown in FIG. 6 described previously. Figure 13

[0379] In the present processing example, it becomes possible to perform control to set a target region to be an output region of a virtual object (i.e., a character) in a substantially horizontal surface such as a ground surface or a floor surface.

[0380] As a result, in a case where a virtual object such as a character is displayed on a real object in the real world, it is possible to display in such a manner that the virtual object comes into contact with a horizontal surface region (e.g., a horizontal surface region on a ground surface in a case where the real object is in a room, or a horizontal surface region on a floor surface in a case where the real object is in a room), and a more natural character display becomes possible, and it becomes possible to make a user feel that the character actually exists in the real world.

[0381] (3 - (3) Update sequence of real object recognition processing)

[0382] Next, an update sequence of real object recognition processing performed by the object recognition unit will be described. ​

[0383] As described previously with reference to Figure 10 , Figure 11 and so on, after the target region is determined, Figure 9 the object recognition unit 126 in the data processing unit 120 of the image processing apparatus 100 shown in Fig. 1 immediately performs recognition processing of the real object in the target region, and thereafter, repeatedly performs the object recognition processing on the region, and sequentially updates the spatial map data shown in Fig. 2. Figure 10

[0384] However, the interval of the update processing varies according to the type (category) of the recognized real object.

[0385] The designation data of the update time which differs according to the type (category) of the real object is registered in advance as the category-associated update time data 128.

[0386] The category-associated update time data 128 is data in which the following data are associated with each other, as described previously with reference to Figure 11 .

[0387] (a) ID

[0388] (b) classification

[0389] (c) category

[0390] (d) update time (seconds)

[0391] The (a) ID is an identifier of the registered data.

[0392] The (b) classification is a classification of the type (category) of the real object.

[0393] The (c) category is the type information of the real object.

[0394] The (d) update time (seconds) is a time indicating the update interval of the real object recognition processing.

[0395] For example, in the case of the category (object type) = lawn of the ID 001, the update time is 3600 seconds (= 1 hour). In the object such as lawn, the change with the passage of time is small, and the update time is set long.

[0396] On the other hand, for example, in the case of the category (object type) = shadow of the ID = 004, the update time is 2 seconds. In the object such as shadow, the change with the passage of time is large, and the update time is set short.

[0397] ​The object recognition unit 126 refers to the data of the class association update time data 128, and repeatedly executes the object recognition processing as necessary at the time interval defined for the recognized object. The real object detected by the new recognition processing is sequentially registered as a reference Figure 10 The spatial map data 127 described above.

[0398] Figure 17 The flowchart illustrated is a flowchart of processing that explains a sequence including repeated execution of the object recognition processing.

[0399] Hereinafter, the processing of each step of the flowchart illustrated will be described. Figure 17 The processing of each step of the flowchart illustrated.

[0400] Note that, Figure 17 The processing of steps S101 to S105 and the processing of steps S106 to S107 of the flowchart illustrated are similar to the processing of each step of the basic processing flow described above with reference to Figure 13 The processing of each step of the basic processing flow described above with reference to

[0401] Figure 17 The processing in steps S401 and S402 of the flow illustrated is a point different from the processing described above. Figure 13 The processing in steps S401 and S402 of the flow illustrated is a point different from the processing described above.

[0402] The processing of each step will be described.

[0403] (Step S401)

[0404] In steps S101 to S105, the determination of the target region and the recognition processing of the real object (class) of the target region are executed, and then the processing of step S401 is executed.

[0405] In step S401, the result of the object recognition processing for the target region executed in step S105 is recorded in the spatial map data.

[0406] As described above with reference to Figure 10 The spatial map data stores the following associated data of each data.

[0407] (a) Time stamp (second)

[0408] (b) Position information

[0409] (c) Class

[0410] (d) Elapsed time (second) after the recognition processing

[0411] (a) Time stamp (second) is time information on the execution of the object recognition processing.

[0412] (b) Position information is position information of the real object that is the object recognition target.

[0413] (c) Category is object type information as a result of object recognition.

[0414] (d) Elapsed time after recognition processing (seconds) is the time that has elapsed since the object recognition processing was completed.

[0415] In step S401 , for the real objects in the target area identified in step S105 , these data are registered in the space map data.

[0416] (Steps S106 to S107)

[0417] The processing of steps S106 to S107 is similar to that of the previous reference Figure 13 That is, the following processing is performed.

[0418] In step S106 , based on the recognized real object (category) in the target area, the type and output mode of the virtual object to be displayed in the target area are determined.

[0419] In step S107 , the virtual object is output (displayed) to the target area according to the type and output mode of the virtual object to be displayed in the target area determined in step S106 .

[0420] (Step S402)

[0421] In addition, after the processing of step S107, in step S402, it is determined whether the time elapsed after the recognition processing of the real object in the target area performed in step S105 exceeds the time in the reference period. Figure 11 The described category is associated with "(d) Update Time" defined in the update time data.

[0422] In the case where it is determined that the elapsed time exceeds the update time, the process returns to step S101 , and the processes of step S101 and subsequent steps are repeatedly performed.

[0423] That is, determination of the target area and real object recognition processing for the target area are performed again.

[0424] In this process, if the position of the target region has not changed, real object recognition is performed again in the target region at the same position.

[0425] On the other hand, if the position of the target region has changed, true object recognition is performed in the target region at the new position.

[0426] By performing these types of processing, it becomes possible to immediately perform the processing of updating the target region and the processing of updating the recognition result of the real object, and a timely virtual object display processing can be performed in accordance with the movement or the instruction of the user.

[0427] [4. Hardware configuration example of image processing apparatus]

[0428] Next, a hardware configuration example of an image processing apparatus that performs the processing according to the above-described embodiments will be described with reference to Figure 18

[0429] Figure 18 The hardware configuration illustrated is an example of the hardware configuration of the image processing apparatus 100 of the present disclosure described above. Figure 9

[0430] A hardware configuration illustrated in Figure 18 will be described.

[0431] A central processing unit (CPU) 301 functions as a data processing unit that performs various types of processing in accordance with a program stored in a read only memory (ROM) 302 or a storage unit 308. For example, processing is performed in accordance with the sequence described in the above-described embodiments. A random access memory (RAM) 303 stores a program, data, and the like executed by the CPU 301. These CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304.

[0432] The CPU 301 is connected to an input / output interface 305 via the bus 304, and the input / output interface 305 is connected to: an input unit 306 including various sensors, a camera, a switch, a keyboard, a mouse, a microphone, and the like; and an output unit 307 including a display, a speaker, and the like.

[0433] The storage unit 308 connected to the input / output interface 305 includes, for example, a hard disk or the like, and stores a program executed by the CPU 301 and various data. A communication unit 309 functions as a data communication transmission / reception unit via a network such as the Internet or a local area network, and further functions as a broadcast wave transmission / reception unit, and communicates with an external apparatus.

[0434] A drive 310 connected to the input / output interface 305 drives a removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory such as a memory card, and performs data recording or reading.

[0435] [5. Summary of configuration of the present disclosure]

[0436] ​​The embodiments of the present disclosure have been described above with reference to specific embodiments. However, it is obvious that a person skilled in the art can modify and replace the embodiments without departing from the gist of the present disclosure. In other words, the present disclosure is disclosed in the form of examples, and should not be construed restrictively. In order to determine the gist of the present disclosure, the scope of the claims should be considered.

[0437] Note that the technology disclosed in this specification can have the following configurations.

[0438] (1) An image processing apparatus comprising:

[0439] an object recognition unit that performs recognition processing of a real object in a real world; and

[0440] a content display control unit that generates an augmented reality (AR) image in which a real object and a virtual object are superimposed and displayed, wherein

[0441] the object recognition unit

[0442] performs object recognition processing of recognizing a real object in a display area of the virtual object, and

[0443] the content display control unit

[0444] selects a virtual object to be displayed in accordance with an object recognition result recognized in the object recognition unit.

[0445] (2) The image processing apparatus according to (1), wherein

[0446] the object recognition unit

[0447] performs object recognition processing by image recognition processing.

[0448] (3) The image processing apparatus according to (2), wherein

[0449] the object recognition unit

[0450] performs object recognition processing by applying semantic segmentation processing.

[0451] (4) The image processing apparatus according to any one of (1) to (3), wherein

[0452] the object recognition unit

[0453] determines a target area to be a display area of the virtual object, and performs recognition processing of a real object in the determined target area.

[0454] (5) The image processing apparatus according to (4), wherein

[0455] the object recognition unit

[0456] determines the target region based on at least any one of a user action, a user line of sight, a user operation, a user position, and a user posture.

[0457] (6) The image processing apparatus according to (4) or (5), wherein

[0458] the object recognition unit

[0459] selects and determines the target region from a horizontal surface region.

[0460] (7) The image processing apparatus according to (6), wherein

[0461] the content display control unit

[0462] displays the virtual object so that the virtual object is in contact with the horizontal surface region.

[0463] (8) The image processing apparatus according to (1) to (7), wherein

[0464] the object recognition unit

[0465] performs the object recognition processing as real-time processing.

[0466] (9) The image processing apparatus according to any one of (1) to (8), wherein

[0467] the object recognition unit

[0468] repeatedly performs the object recognition processing at a predefined time interval according to an object type.

[0469] (10) The image processing apparatus according to any one of (1) to (9), further comprising:

[0470] a three-dimensional map generation unit that generates a three-dimensional map of a real world based on an image captured by the imaging device, wherein

[0471] the object recognition unit

[0472] determines a target region to be a display region of the virtual object by using the three-dimensional map.

[0473] (11) The image processing apparatus according to (10), wherein

[0474] the three-dimensional map generation unit

[0475] generates the three-dimensional map of the real world by simultaneous localization and mapping (SLAM) processing.

[0476] (12) The image processing apparatus according to any one of (1) to (11), in which

[0477] the content display control unit

[0478] selects a virtual object to be displayed in accordance with an object recognition result recognized in the object recognition unit, and

[0479] controls a display mode of the virtual object to be displayed also in accordance with the object recognition result.

[0480] (13) The image processing apparatus according to any one of (1) to (12), further comprising:

[0481] a content sound control unit that performs sound output control, in which

[0482] the content sound control unit

[0483] determines and outputs a sound to be output in accordance with an object recognition result recognized in the object recognition unit.

[0484] (14) An image processing method executed in an image processing apparatus, the method comprising:

[0485] an object recognition processing step performed by an object recognition unit that performs recognition processing of a real object in a real world; and

[0486] a content display control step performed by a content display control unit that generates an augmented reality (AR) image in which a real object and a virtual object are superimposed and displayed, in which

[0487] the object recognition processing step

[0488] is a step that performs object recognition processing that recognizes a real object in a display area of the virtual object, and

[0489] the content display control step

[0490] performs a step of selecting a virtual object to be displayed in accordance with an object recognition result recognized in the object recognition processing step.

[0491] (15) A program for causing an image processing to be executed in an image processing apparatus, the program:

[0492] causes an object recognition unit to perform an object recognition processing step that performs recognition processing of a real object in a real world;

[0493] The content display control unit is caused to execute a content display control step of generating an augmented reality (AR) image in which a real object and a virtual object are superimposed and displayed;

[0494] In the object recognition processing step,

[0495] The object recognition processing is caused to be executed, the object recognition processing recognizing a real object in a display area of the virtual object; and

[0496] In the content display control step,

[0497] The step of selecting a virtual object to be displayed according to the object recognition result recognized in the object recognition processing step is caused to be executed.

[0498] Further, a series of processing steps described in the specification can be executed by hardware, software, or a combination of both. In the case where the processing is executed by software, a program that records the processing sequence in a memory can be installed and executed in a computer included in a dedicated hardware, or in a general-purpose computer capable of executing various different types of functions. For example, the program can be recorded in a recording medium in advance. In addition to being installed from the recording medium to the computer, the program can be received via a network such as a local area network (LAN) or the Internet, and installed in a recording medium such as a built-in hard disk.

[0499] Note that the various types of processing described in the specification are not only executed in time series according to the specification, but can also be executed in parallel or individually according to the processing capacity of the device that executes the processing or as needed. Further, in the present specification, the term "system" is a logical group configuration of a plurality of devices, and is not limited to the configuration of each configured device in the same housing.

[0500] Industrial applicability

[0501] As described above, according to the configuration of the embodiment of the present disclosure, a device and a method that perform selection or display mode change of a virtual object to be displayed according to a real object type in a target area to be a display area of the virtual object are realized.

[0502] Specifically, for example, including: an object recognition unit that executes recognition processing of a real object in a real world, and a content display control unit that generates an AR image in which a real object and a virtual object are superimposed and displayed. The object recognition unit recognizes a real object in a target area to be a display area of a virtual object, and the content display control unit executes processing of selecting a virtual object to be displayed or processing of changing a display mode according to the object recognition result.

[0503] With this configuration, an apparatus and method are realized that perform selection or display mode change of a virtual object to be displayed according to a real object type in a target region to be a display region of the virtual object.

[0504] List of reference numerals

[0505] 10 Light-transmissive AR image display device

[0506] 11 Target region

[0507] 21 Transmissive observation image

[0508] 22 to 24 Virtual object image

[0509] 30 Image display type AR image display device captured by image pickup device

[0510] 31 Image pickup device

[0511] 32 Image captured by image pickup device

[0512] 40 Smartphone

[0513] 41 Image pickup device

[0514] 42 Image captured by image pickup device

[0515] 50 Virtual object image on water

[0516] 51 Virtual object image in water

[0517] 52 Virtual object image

[0518] 53, 54 Virtual object shadow image

[0519] 100 Image processing apparatus

[0520] 110 Data input unit

[0521] 111 External imaging image pickup device

[0522] 112 Internal imaging image pickup device

[0523] 113 Motion sensor (gyro, acceleration sensor, etc.)

[0524] 114 Operation unit

[0525] 115 Microphone

[0526] 120 Data processing unit

[0527] 121 External captured image analysis unit

[0528] 122 Three-dimensional map generation unit

[0529] 123 internal captured image analysis unit

[0530] 124 device posture analysis unit

[0531] 125 sound analysis unit

[0532] 126 object recognition unit

[0533] 127 spatial map data

[0534] 128 category association update time data

[0535] 130 data output unit

[0536] 131 content display control unit

[0537] 132 content sound control unit

[0538] 133 display unit

[0539] 134 speaker

[0540] 135 category association virtual object data (3D model, sound data, etc.)

[0541] 140 communication unit

[0542] 301 CPU

[0543] 302 ROM

[0544] 303 RAM

[0545] 304 bus

[0546] 305 input / output interface

[0547] 306 input unit

[0548] 307 output unit

[0549] 308 storage unit

[0550] 309 communication unit

[0551] 310 driver

[0552] 311 removable medium

Claims

1. An information processing system comprising: a first information processing device configured to receive user operation information; as well as A head-mounted display comprising circuitry configured to: receiving an image of a real space captured by an image capture device; identifying at least one feature point based on the received image; generating a three-dimensional map of the real space based on the identified at least one feature point; determining a first target area based on the generated three-dimensional map; as well as controlling the display to display a first virtual object, wherein the first virtual object is related to a real object in the first target area, and Wherein, the first information processing device is spatially separated from the head-mounted display.

2. The information processing system according to claim 1, wherein The first virtual object is displayed in spatial relation to the first target area based on information related to the first information processing device.

3. The information processing system according to claim 2, wherein: The information related to the first information processing device includes direction information indicating a direction of the first information processing device.

4. The information processing system according to claim 1, wherein: The circuit is further configured to determine a second target area based on the user operation information input from the first information processing device.

5. The information processing system according to claim 4, wherein: The circuit is further configured to control the display to display a second virtual object in spatial relation to the second target area, wherein the second virtual object is different from the first virtual object, and The second virtual object is related to the real object in the second target area. The information processing system according to claim 5 , wherein: The second virtual object is displayed in spatial relation to the second target area based on information related to the first information processing device.

7. The information processing system according to claim 1, wherein: The first information processing device includes a bar-shaped member.

8. An information processing method, comprising: receiving user operation information from the first information processing device; receiving an image of a real space captured by an image capture device; identifying at least one feature point based on the received image; generating a three-dimensional map of the real space based on the identified at least one feature point; determining a first target area based on the generated three-dimensional map; as well as displaying a first virtual object, The first virtual object is related to a real object in the first target area.

9. The information processing method according to claim 8, wherein: The first virtual object is displayed in spatial relation to the first target area based on information related to the first information processing device.

10. The information processing method according to claim 9, wherein: The information related to the first information processing device includes direction information indicating a direction of the first information processing device.

11. The information processing method according to claim 8, further comprising: A second target area is determined based on the user operation information input from the first information processing apparatus.

12. The information processing method according to claim 11, further comprising: displaying a second virtual object in spatial relation to the second target area, wherein the second virtual object is different from the first virtual object, and The second virtual object is related to the real object in the second target area.

13. The information processing method according to claim 12, wherein: The second virtual object is displayed in spatial relation to the second target area based on information related to the first information processing device.

14. A non-transitory computer-readable storage medium having a program embodied thereon, wherein when the program is executed by a computer, the computer is caused to perform an information processing method, the information processing method comprising: receiving user operation information from the first information processing device; receiving an image of a real space captured by an image capture device; identifying at least one feature point based on the received image; generating a three-dimensional map of the real space based on the identified at least one feature point; determining a first target area based on the generated three-dimensional map; as well as displaying a first virtual object, The first virtual object is related to a real object in the first target area.

15. The non-transitory computer-readable storage medium of claim 14, wherein: The first virtual object is displayed in spatial relation to the first target area based on information related to the first information processing device.

16. The non-transitory computer-readable storage medium of claim 15, wherein: The information related to the first information processing device includes direction information indicating a direction of the first information processing device.

17. The non-transitory computer-readable storage medium of claim 14, wherein: The information processing method further includes: A second target area is determined based on the user operation information input from the first information processing apparatus.

18. The non-transitory computer-readable storage medium of claim 17, wherein: The information processing method further includes: displaying a second virtual object in spatial relation to the second target area, wherein the second virtual object is different from the first virtual object, and The second virtual object is related to the real object in the second target area.

19. The non-transitory computer-readable storage medium of claim 18, wherein: The second virtual object is displayed in spatial relation to the second target area based on information related to the first information processing device.

Citation Information

Patent Citations

  • Information processing system, information processing method, and information processing program

    JP2015060579A

  • Head-mounted type display device, method of controlling head-mounted type display device, and computer program

    JP2017091433A

  • Methods and systems for creating virtual and augmented reality

    US20190094981A1

  • Virtual user input controls in a mixed reality environment

    WO2018106542A1

  • Image recognition device, image recognition method, and program

    WO2019016870A1