Ar control method
The AR control method improves AR effect generation accuracy by using a mask image with specific RGB values to differentiate active and inactive areas, addressing misrecognition issues in AR detection.
Patent Information
- Application Number
- PCT/JP2025/023787
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-08
AI Technical Summary
Existing AR control methods face challenges in accurately detecting and generating corresponding AR effects due to similarities between different states, leading to misrecognition and incorrect effect generation.
An AR control method that utilizes a mask image with specific RGB values to distinguish active and inactive areas, allowing for precise detection and generation of AR effects based on the placement of partial markers within a defined shape, reducing erroneous extractions by focusing on effective areas.
Enhances the accuracy of AR effect generation by minimizing misrecognition through precise detection of marker placement, ensuring correct AR effects are generated based on the state of the marker arrangement.
Smart Images

Figure JP2025023787_08012026_PF_FP_ABST
Abstract
Description
AR control method
[0001] The present invention relates to an AR control method.
[0002] An AR control method is known that detects an AR marker and generates a corresponding AR effect.
[0003] A new AR control method is provided.
[0004] The following configuration is provided as an example.
[0005] [1] An AR control method in which: - an image of an overall marker; - characteristics of first to n-th states when a state k is a state in which first to k-th partial markers are respectively arranged in first to k-th (k=1 to n, n is an integer equal to or greater than 2) regions of a shape corresponding to the overall marker; - first to n-th AR effects associated with each of the first to n-th states; and the method comprises: a first step of recognizing an area marker of the shape; a second step of recognizing the area marker of the image; a third step of confirming an area marker in a correct position in the area; a fourth step of determining the current overall marker as a mask; a fifth step of overlapping the mask and the overall marker; and a sixth step of generating an AR effect based on the result of the fifth step.
[0006] [2] The AR control method according to [1], wherein the mask is a mask image in which the RGB values of an inactive area are R, G, B = 0, 0, 0 and the RGB values of an active area are R, G, B = 1, 1, 1, and the fifth step includes multiplying the mask image by the whole image marker and executing the result.
[0007] [3] A method for identifying marker objects in an augmented reality environment, comprising the steps of: providing an overall image stored in a database; reading the image; creating a transparency channel layer for the image, setting black areas of the image to transparent and colored areas of the image to opaque; multiplying the image by the transparency channel layer to obtain a target image for comparison; and performing a color comparison between the target image and corresponding areas of the overall image to obtain a comparison result.
[0008] [4] The method according to [3], wherein the transparency channel layer is multiplied with the overall image to obtain a comparison region image, and a color comparison is performed between the target image and the comparison region image to obtain a comparison result.
[0009] [5] The method according to [3], further comprising a normalization step, i.e., a step of normalizing the values of the transparency channel layer to a range from 0 to 1.
[0010] [6] The method of [5], wherein the normalizing step includes adjusting the level of the transparency channel layer to normalize values to 0 or 1.
[0011] [7] The method according to [3], further comprising a resizing step, i.e., a step of resizing the size of the image to fit the overall size in order to optimize the comparison result.
[0012] [8] The method according to [3], wherein the overall image is divided into at least one first region, and the at least one first region includes regional parameters, the method comprising: performing a color comparison between the target image and the overall image in the at least one first region, and if the comparison results in a match, recording the regional parameters of the at least one first region; and performing a feedback response based on the recorded regional parameters.
[0013] [9] The method of [8], wherein the at least one first region comprises a plurality of first regions, and the feedback response is based on a sum of the recorded region parameters.
[0014] 1 is a diagram for explaining an overview of the present invention. FIG. ... example of a puzzle assumed in this embodiment. A block diagram showing a general configuration of an AR control device according to an embodiment. A diagram schematically showing data stored in a storage unit 4. A flowchart showing an example of the processing operation of the AR control device. A diagram for explaining the processing operation of the AR control device in state A where no partial markers have yet been placed. A diagram for explaining the processing operation of the AR control device when a user places an ear partial marker to change from state A to state B. A diagram for explaining the processing operation of the AR control device when state C where ear and face partial markers have been placed is presented.
[0015] The present invention is outlined in Figures 1A to 1I, which will be described in detail below.
[0016] In this embodiment, it is assumed that an AR (Augmented Reality) effect is generated according to the degree of completion of the puzzle. A specific example is shown in FIG. 2. The illustrated puzzle involves collecting multiple partial markers (pieces) and placing them in a predetermined area on a backing to complete the entire marker. The AR effect is, for example, superimposing a predetermined image on the real space captured by a camera on the display of a user terminal.
[0017] FIG. 2(a) shows a state where no partial markers have been collected (hereinafter referred to as "State A"). In State A, a black shape corresponding to the overall marker (bear) is displayed on the mount. The overall marker is formed from multiple areas corresponding to each of the partial markers described below. When such a mount is photographed by the camera of the user terminal, an AR effect corresponding to State A (hereinafter referred to as "AR Effect A") is generated on the display of the user terminal (or no AR effect is generated).
[0018] FIG. 2(b) shows a state (hereinafter referred to as "state B") in which the user progresses through the puzzle and collects one of the partial markers (the ear). That is, on the backing sheet, the ear partial marker is placed in the area of the black shape that corresponds to the ear. When the backing sheet with such partial markers placed thereon is photographed by the camera of the user terminal, an AR effect corresponding to state B (hereinafter referred to as "AR effect B") is generated on the display of the user terminal.
[0019] 2(c) shows a state (hereinafter referred to as "state C") in which the user has progressed further in the puzzle and collected two of the partial markers (the ear and the face). That is, on the backing sheet, the partial markers for the ear and the face are arranged in the areas of the black shape corresponding to the ear and the face, respectively. When the backing sheet on which such partial markers are arranged is photographed by the camera of the user terminal, an AR effect corresponding to state C (hereinafter referred to as "AR effect C") is generated on the display of the user terminal.
[0020] 2(d) shows a state (hereinafter referred to as "state D") in which the user has progressed further in the puzzle and collected three of the partial markers (ear, face, and right foot). That is, on the backing sheet, partial markers for the ear, face, and right foot are arranged in areas of the black shape corresponding to the ear, face, and right foot, respectively. When the backing sheet on which such partial markers are arranged is photographed by the camera of the user terminal, an AR effect corresponding to state D (hereinafter referred to as "AR effect D") is generated on the display of the user terminal.
[0021] 2(e) shows a state (hereinafter referred to as "state E") in which the user has progressed further in the puzzle and collected four of the partial markers (ear, face, right foot, and left foot). That is, on the backing sheet, partial markers for the ear, face, right foot, and left foot are arranged in areas of the black shapes corresponding to the ear, face, right foot, and left foot, respectively. When the backing sheet on which such partial markers are arranged is photographed by the camera of the user terminal, an AR effect corresponding to state E (hereinafter referred to as "AR effect E") is generated on the display of the user terminal.
[0022] 2(f) shows the state in which the user has progressed further in the puzzle and completed it, in other words, the state in which five of the partial markers (ears, face, right foot, left foot, and torso) have been collected (hereinafter referred to as "state F"). That is, on the backing paper, partial markers for the ears, face, right foot, left foot, and torso are arranged in areas of the black shapes corresponding to the ears, face, right foot, left foot, and torso, respectively. When the backing paper on which such partial markers are arranged is photographed by the camera of the user terminal, an AR effect corresponding to state F (hereinafter referred to as "AR effect F") is generated on the display of the user terminal.
[0023] Here, one possible method for the system to grasp the progress of the puzzle by taking a picture with a camera is to compare the entire image captured by the camera with the images of each state shown in Figures 2(a) to 2(f). The image of state B and the image of state C are significantly different, making them unlikely to be misrecognized. On the other hand, the image of state E and the image of state F are similar, making them likely to be misrecognized. In that case, AR effect F may be generated even though the player is in state E, or AR effect E and AR effect F may be generated alternately.
[0024] 3 is a block diagram showing a schematic configuration of an AR control device according to an embodiment. The AR control device includes a camera 1, a database 2, a system 3, and the like.
[0025] As shown in Fig. 4(c), AR effects A to F associated with states A to F are stored, respectively. Note that if no AR effect is generated in state A where no partial markers are placed, there is no AR effect A corresponding to state A.
[0026] 5 is a flowchart showing an example of the processing operation of the system. First, the camera 1 captures an image of a mount on which one or more partial markers are arranged (or not arranged) (step S1).
[0027] Next, the system 3 recognizes the shape (outline) corresponding to the whole marker from the database 2 (step S2). Here, the system 3 only needs to recognize the shape. After the shape detection is determined, the system 3 recognizes the image. In this procedure, if the shape does not match the database 3, the operation stops. Similarly, if the image does not match the database 3, the operation stops. Based on this method, the amount of calculation of the system is reduced and the efficiency is improved.
[0028] Then, the system 3 recognizes the marker placed in the area of the mount and determines the current mount as the reference marker.
[0029] Next, the system 3 creates a mask based on the reference markers (step S3). Black areas are marked as R, G, B = 0, 0, 0, and determined as one piece of information, and other areas are marked as R, G, B = 1, 1, 1, and determined as another piece of information.
[0030] Next, System 3 compares the mask with the global markers from Database 2. System 3 multiplies the mask with the global (complete) image and obtains the following results: - The mask area has R,G,B=0,0,0, and the other areas of the mask have R,G,B=1,1,1. And the global image from Database 2 has R,G,B=1,1,1. After multiplication, the result is: 1 The mask area has R,G,B=0,0,0 x R,G,B=1,1,1, resulting in R,G,B=0,0,0 2 The other areas of the mask have R,G,B=1,1,1 x R,G,B=1,1,1, resulting in R,G,B=1,1,1
[0031] Then, the system 3 determines the result of R, G, B=1, 1, 1 and generates the corresponding AR effect.
[0032] Since the detection is based on the characteristics of the effective area, rather than on the detection of each marker, erroneous extraction can be reduced. Also, since the mask image is created based on RGB information, the system only needs to judge a small amount of information and does not need to read a large amount of information, so the AR effect is generated with high accuracy.
[0033] Several scenarios will be specifically described below: Figure 6 is a diagram illustrating the processing operation of the system in a state A where no partial markers have been placed yet.
[0034] As shown in Fig. 6(a), a user presents a mount on which no partial markers are placed, and the mount is photographed by camera 1 (step S1). Then, the system recognizes the shape corresponding to the whole marker from the captured image (step S2).
[0035] As shown in Fig. 6(a), the system recognizes that the entire shape corresponding to the entire marker is black (step S4). Therefore, as shown in Fig. 6(b), a mask image is created in which the RGB information of the entire area of the shape corresponding to the entire marker is set to R, G, B = 0, 0, 0 (step S3).
[0036] Then, as shown in Figure 6(c), the system multiplies the image of the overall marker by the mask image (step S5). Because the RGB values of the entire area of the mask image are 0, the RGB values of the area corresponding to the overall marker in the multiplication result are all 0, and there is no valid area. Therefore, there is no corresponding state, and none of states B to F are extracted (step S6). Therefore, no AR effect is generated (step S7).
[0037] FIG. 7 is a diagram for explaining the processing operation of the system when the user places the partial marker on the ear to change from state A to state B.
[0038] 7A shows a state in which the user is attempting to place a partial ear marker but has not yet completed placement. In this case, the shape of the entire marker is recognized from the image captured by camera 1. However, because this shape is entirely black, no AR effect is generated, as in FIG. 6.
[0039] 7(b) shows a state in which the user has placed the ear partial marker in an appropriate area of the overall marker on the mount. Then, the system recognizes the shape corresponding to the overall marker from the acquired image (step S2).
[0040] As shown in Fig. 7(b), the system recognizes that the shape corresponding to the whole marker is black except for the ear region (step S4). Therefore, as shown in Fig. 7(c), the system creates a mask image in which the R, G, and B values of the shape corresponding to the whole marker are all 0 in the region other than the ear, and all R, G, and B values of the ear region are 1 (step S4).
[0041] 7(d), the system multiplies the image of the overall marker by the mask image (step S5). As described above, in the mask image, the RGB values of the areas other than the ear area are 0, and the pixel value of the ear area is 1. Therefore, of the areas corresponding to the overall marker in the multiplication result, only the ear area has an RGB value that is not 0, and is therefore a valid area.
[0042] Therefore, state B, the characteristics of which match this effective area, is extracted from the storage unit 4 (step S6), and AR effect B is generated (step S7).
[0043] FIG. 8 is a diagram illustrating the processing operation of the system when state C in which partial markers for the ear and face have been placed is presented.
[0044] 8(a) shows a state in which the user has already placed ear and face markers in appropriate areas of the overall marker on the mount. The system then recognizes the shape corresponding to the overall marker from the acquired image (step S2).
[0045] 8(a), the system recognizes that the shape corresponding to the whole marker is black except for the ear and face regions (step S4). Therefore, as shown in FIG. 8(b), the system creates a mask image in which the R, G, and B values of the shape corresponding to the whole marker are all 0 in the regions other than the ear and face, and all R, G, and B values of the ear and face regions are 1 (step S4).
[0046] 8(c), the system multiplies the image of the overall marker by the mask image (step S5). As described above, in the mask image, the RGB values of the areas other than the ears and face are 0, and the RGB values of the ears and face are 1. Therefore, of the areas corresponding to the overall marker in the multiplication result, only the ears and face areas have non-zero values and are valid areas.
[0047] Therefore, a state C whose characteristics match this effective region is extracted from the database 2 (step S6), and an AR effect C is generated (step S7).
[0048] As described above, this system recognizes areas with RGB information from the image captured by camera 1 and creates a mask. This mask is used to extract the effective area from the overall marker. This allows for highly accurate extraction of the state and the generation of a correct AR effect.
[0049] This embodiment can be applied to puzzles in museums. For example, a mount is provided to museum visitors. Then, when the visitors view a certain exhibition zone, they receive a partial marker. When the partial marker is placed on the mount, an AR effect is generated to guide the visitors to the next exhibition zone. When all exhibition zones are viewed, all partial markers are placed on the mount, and an AR effect related to the museum is generated.
[0050] This embodiment can be applied to a game including multiple missions (e.g., an escape room). For example, when a game user completes one mission, they receive a partial marker. When the partial marker is placed on the mount, an AR effect for the next mission is generated. This embodiment can also be applied to various other games, such as integrated games featuring anime characters.
[0051] Any part or all of the functional units described in this specification may be realized by a program. The program mentioned in this specification may be non-transitoryly recorded on a computer-readable recording medium.
[0052] Such a program may be installed on a computer (a so-called native application), in which case the program may be downloaded to the computer via a communication line (including wireless communication) such as the Internet, or may be distributed in a state where it is installed on the computer.
[0053] Based on the above description, a person skilled in the art may be able to conceive additional effects and various modifications of the present invention, but the aspects of the present invention are not limited to the individual embodiments described above. For example, inventions that extract only a part of each embodiment or inventions that combine multiple embodiments are naturally envisioned. Various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and spirit of the present invention, which can be derived from the content defined in the claims and their equivalents.
[0054] For example, what is described in this specification as one device (or component, the same applies hereinafter) (including what is depicted as one device in the drawings) may be realized by multiple devices. Conversely, what is described in this specification as multiple devices (including what is depicted as multiple devices in the drawings) may be realized by one device. Alternatively, some or all of the means or functions included in one device may be included in another device. Furthermore, a "system" may be composed of one device, or two or more devices.
[0055] Furthermore, not all of the features described in this specification are essential requirements. In particular, features described in this specification but not included in the claims can be considered optional additional features.
[0056] Furthermore, unless otherwise specified, the term "means" in this specification and claims refers to hardware itself (or a function realized by hardware) and does not include a human being (or human mental activity).
[0057] 1. Camera 2. Database 3. System
Claims
1. An AR control method in which: - an image of an overall marker; - characteristics of first to nth states when the kth state is a state in which first to kth partial markers are placed in first to kth (k = 1 to n, n is an integer of 2 or more) regions of a shape corresponding to the overall marker; - first to nth AR effects associated with each of the first to nth states; and the method comprises: a first step of recognizing an area marker of the shape; a second step of recognizing the area marker of the image; a third step of confirming an area marker in a correct position in the region; a fourth step of determining the current overall marker as a mask; a fifth step of overlapping the mask and the overall marker; and a sixth step of generating an AR effect based on the result of the fifth step.
2. The AR control method of claim 1, wherein the mask is a mask image in which the RGB values of the inactive area are R, G, B = 0, 0, 0 and the RGB values of the active area are R, G, B = 1, 1, 1, and the fifth step includes multiplying the mask image by the whole image marker and executing the result.
3. A method for identifying marker objects in an augmented reality environment, comprising the steps of: providing an overall image stored in a database; reading the image; creating a transparency channel layer for the image, setting black areas of the image to transparent and colored areas of the image to opaque; multiplying the image by the transparency channel layer to obtain a target image for comparison; and performing a color comparison between the target image and corresponding areas of the overall image to obtain a comparison result.
4. The method of claim 3, wherein the transparency channel layer is multiplied with the overall image to obtain a comparison region image; and a color comparison is performed between the target image and the comparison region image to obtain a comparison result.
5. The method of claim 3, further comprising a normalization step, i.e., a step of normalizing the values of the transparency channel layer to the range 0 to 1.
6. The method of claim 5, wherein said normalizing step includes adjusting the level of said transparency channel layer to normalize values to 0 or 1.
7. The method of claim 3, further comprising a resizing step, i.e., a step of resizing the size of the image to fit the overall size in order to optimize the comparison result.
8. The method of claim 3, wherein the overall image is divided into at least one first region, the at least one first region including regional parameters, the method comprising: performing a color comparison between the target image and the overall image in the at least one first region, and if the comparison results in a match, recording the regional parameters of the at least one first region; and performing a feedback response based on the recorded regional parameters.
9. The method of claim 8, wherein the at least one first region comprises a plurality of first regions, and the feedback response is based on a sum of the recorded region parameters.
Citation Information
Patent Citations
Method for generating personalized product views
EP2717226A1
Card learning system and card learning method
JP2021060440A
Method for recognizing object assemblies in augmented reality images
JP2024077001A