Interactive projection method and device and storage medium
By building a three-dimensional interaction system that connects vision, hearing and behavior in the projection device, the problems of single functions and insufficient interactive capabilities of the existing projection device are solved, and the immersive experience and fun in multiple user scenarios are improved.
Patent Information
- Application Number
- CN202510183459.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-11
AI Technical Summary
The existing projection equipment has a single function, lacks multi-person interaction ability, and has poor immersive experience.
In the preset interactive activities involving multiple users, the target object in the target area is determined, and the virtual image is projected to the target object based on the interaction model, user reaction information is obtained, the guide speech is determined and the beam is projected, and a three-dimensional interaction system that links vision, hearing and behavior is constructed.
显著提高了多用户场景下的沉浸式参与体验和交互活动的趣味性和协作性,实现了更具沉浸感的游戏娱乐体验。
Smart Images

Figure CN120295597A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of projection technology, and in particular, to an interactive projection method, device, equipment, and storage medium. Background Art
[0002] Projection devices are a type of multimedia device that are widely used in daily life. And with the continuous development of technology, their application scenarios are also becoming more and more extensive, such as in homes, offices, schools, and entertainment venues.
[0003] However, the existing projection devices are still relatively traditional, with single functions and single structures. During the process of organizing meetings or games, they do not have the ability to interact with multiple people, and the immersive experience effect is relatively poor. Summary of the Invention
[0004] The present disclosure provides an interactive projection method, equipment, and storage medium to at least solve the above technical problems existing in the prior art.
[0005] According to a first aspect of the present disclosure, an interactive projection method is provided. The method includes:
[0006] During the process of completing a preset interaction activity participated by multiple users, determine the target objects within the target area. Among them, the target objects include first target objects, and each first target object corresponds to a user one by one;
[0007] Based on the interaction model and the first target objects, project the corresponding guiding virtual images to each first target object, where the interaction model is a model trained based on the content involved in the preset interaction activity;
[0008] Within the target area, obtain the reaction information made by each user based on the guiding virtual images;
[0009] Based on the interaction model and the reaction information, determine the first guiding voice, where the first guiding voice is used to promote the preset interaction activity;
[0010] Determine the first target user through the first guiding voice, and project a first light beam to the first target user.
[0011] In an implementable manner, after projecting the first light beam to the first target user, it further includes:
[0012] Real-time obtain the first reaction information of the first target user;
[0013] When determining that the first reaction information is an end information, determine the second target user based on the interaction model;
[0014] Move the first light beam to the second target user and output a second guiding voice, where the content of the second guiding voice is different from that of the first guiding voice.
[0015] In one implementable manner, determining the object within the target area includes:
[0016] Obtain the original image within the target area, where the original image includes multiple objects;
[0017] Identify the original image through an identification model to obtain the two-dimensional image coordinates of the multiple objects, where the identification model is used to identify the multiple objects;
[0018] Obtain the depth image of the object within the target area;
[0019] Determine the true three-dimensional coordinates of the contour points of the multiple objects through the depth image and the two-dimensional image coordinates of the multiple objects;
[0020] Determine each object and its corresponding target label according to the true three-dimensional coordinates of each object contour point, where the target label is used to determine the guiding virtual image to be projected by each object by comparing with the interaction model.
[0021] In one implementable manner, projecting the guiding virtual image corresponding to each first object to each first object based on the interaction model includes:
[0022] Restore each object in the virtual space based on the true three-dimensional coordinates of each object contour point;
[0023] Determine the guiding virtual image to be projected by the first object based on the target label and the interaction model;
[0024] Project the guiding virtual image in the peripheral area of the first object.
[0025] In one implementable manner, when determining that the first response information is the end information, determining the second target user based on the interaction model includes:
[0026] When the first response information is the voice information of the first target user, determine the semantic information of the voice information through a semantic model;
[0027] Based on the semantic information, determine whether the first target user terminates speaking; and,
[0028] Determine the second target user based on the interaction model and the semantic information.
[0029] In one possible implementation, it further includes:
[0030] When the voice information of the third target user is collected, a second light beam is projected towards the third target user.
[0031] In one possible implementation, determining the first guiding voice based on the interaction model and the response information includes:
[0032] Based on the interaction model, determine the node advancement condition;
[0033] When the response information meets the node advancement condition, determine the first guiding voice.
[0034] In one possible implementation, wherein the response information includes facial information, gesture information, and / or voice information, and the response information meeting the node advancement condition includes at least one of the following:
[0035] Obtain the facial information of each user, when there is at least one user's facial information that meets the facial advancement condition;
[0036] Obtain the gesture information of each user, when there is at least one user's gesture information that meets the gesture advancement condition;
[0037] Obtain the voice information of each user, when there is at least one user's voice information that meets the voice advancement condition.
[0038] According to the second aspect of the present disclosure, an interactive projection device is provided, and the device includes:
[0039] A target determination module, configured to determine a target within a target area range during the process of completing a preset interaction activity participated by multiple users, wherein the target includes a first target, and the first target corresponds to a user one by one;
[0040] An image projection module, configured to project the corresponding guiding virtual image to each first target based on the interaction model and the first target, wherein the interaction model is a model trained based on the content involved in the preset interaction activity;
[0041] An acquisition module, configured to acquire the response information made by each user based on the guiding virtual image within the target area range;
[0042] A voice determination module, configured to determine a first guiding voice based on the interaction model and the response information, wherein the first guiding voice is used to advance the preset interaction activity;
[0043] A light beam projection module, configured to determine a first target user based on the first guiding voice and project a first light beam towards the first target user.
[0044] In an implementable embodiment, it further includes:
[0045] A light beam movement module, configured to, after projecting the first light beam towards the first target user, acquire first response information of the first target user in real time; when determining that the first response information is end information, determine a second target user based on the interaction model; move the first light beam to the second target user, and output a second guiding voice, where the content of the second guiding voice is different from that of the first guiding voice.
[0046] In an implementable embodiment, the target object determination module is specifically configured to: acquire an original image within the target area, where the original image includes multiple target objects; identify the original image through an identification model to obtain two-dimensional image coordinates of the multiple target objects, where the identification model is used to identify the multiple target objects; acquire a depth image of the target objects within the target area; determine real three-dimensional coordinates of contour points of the multiple target objects through the depth image and the two-dimensional image coordinates of the multiple target objects; determine each target object and its corresponding target label according to the real three-dimensional coordinates of each target object contour point, where the target label is used to determine a guiding virtual image to be projected on each target object by comparing with the interaction model.
[0047] In an implementable embodiment, the image projection module is specifically configured to: restore each target object in a virtual space based on the real three-dimensional coordinates of each target object contour point; determine a guiding virtual image to be projected on the first target object based on the target label and the interaction model; project the guiding virtual image on a peripheral area of the first target object.
[0048] In an implementable embodiment, the light beam movement module is specifically configured to: when the first response information is voice information of the first target user, determine semantic information of the voice information through a semantic model; based on the semantic information, determine whether the first target user terminates speaking; and determine the second target user based on the interaction model and the semantic information.
[0049] In an implementable embodiment, it further includes:
[0050] A projection module, configured to project a second light beam towards a third target user when voice information of the third target user is collected.
[0051] In one implementable manner, the voice determination module is specifically configured to: determine a node advancement condition based on an interaction model; and when the response information meets the node advancement condition, determine the first guiding voice.
[0052] In one implementable manner, the response information meeting the node advancement condition includes at least one of the following: obtaining the facial information of each user, when there is at least one user's facial information that meets the facial advancement condition; obtaining the gesture information of each user, when there is at least one user's gesture information that meets the gesture advancement condition; obtaining the voice information of each user, when there is at least one user's voice information that meets the voice advancement condition.
[0053] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0054] at least one processor; and
[0055] a memory communicatively connected to the at least one processor; wherein,
[0056] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the present disclosure.
[0057] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method described in the present disclosure.
[0058] In the interactive projection method, device and storage medium of the present disclosure, during the process of completing a preset interaction activity participated by multiple users, a target object within a target area is determined, and based on the interaction model and the first target object, a corresponding guiding virtual image is projected onto each first target object; within the target area, response information made by each user based on the guiding virtual image is obtained; based on the interaction model and the response information, a first guiding voice is determined; a first target user is determined through the first guiding voice, and a first light beam is projected onto the first target user. By constructing a three-dimensional interaction system with visual, auditory, and behavioral linkages, the immersive participation experience in a multi-user scenario can be effectively improved, and the interestingness and cooperation of the interaction activity can be significantly enhanced.
[0059] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings
[0060] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, where:
[0061] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0062] Figure 1 A schematic diagram showing the implementation process of an interactive projection method provided by an embodiment of the present disclosure is shown;
[0063] Figure 2 A schematic diagram showing an exemplary guiding virtual image projection provided by an embodiment of the present disclosure is shown;
[0064] Figure 3 Another schematic diagram showing an exemplary guiding virtual image projection provided by an embodiment of the present disclosure is shown;
[0065] Figure 4 A schematic diagram showing the refined process of step S110 provided by an embodiment of the present disclosure is shown;
[0066] Figure 5 A schematic diagram showing the structure of an interactive projection device provided by an embodiment of the present disclosure is shown;
[0067] Figure 6 A schematic diagram showing the composition structure of an electronic device according to an embodiment of the present disclosure is shown. Detailed Embodiments
[0068] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present disclosure.
[0069] Figure 1 A flowchart of an interactive projection method provided by an embodiment of the present disclosure. This method can be executed by the interactive projection device provided by an embodiment of the present disclosure, and the device can be implemented in a software and / or hardware manner. The method specifically includes:
[0070] S110. During the process of completing a preset interaction activity participated by multiple users, determine the target object within the target area range.
[0071] Wherein, the target object includes a first target object, and the first target object corresponds to each user one by one.
[0072] Among them, the preset interaction activity can be, for example, an office meeting or a party game participated by multiple people. The target area range can be the range where the user conducts activities during the preset interaction activity, and it is also the area where images need to be collected for activity recognition, so as to track the activity process, interpret the activity, and normally promote the preset interaction activity.
[0073] Among them, the target object can be some supplies required to ensure the smooth progress of the preset interaction activity. Exemplarily, if the preset interaction activity is a Werewolf game in an offline board game scenario, the target object can be prop cards, batons, number lights, masks, etc. used during the game. Among them, the target object can include a first target object and a second target object. The first target object is the target object that needs to be interpreted or given guiding prompts, and the second target object is the target object that does not need to be interpreted or given guiding prompts. And, the first target object corresponds to each user one by one.
[0074] Exemplarily, if the preset interaction activity is a Werewolf game in an offline board game scenario, the target area range can be the tabletop, the first target object can be a card, and each user has a first target object (a card) corresponding to them, and the contents of different first target objects are different. For example, the contents of different identity cards are different; the second target object can be a baton. Exemplarily, in this embodiment, at the beginning of the game content, the user can realize the registration of their respective [domains] by placing their palms on the tabletop, and the gesture instructions and voice instructions within the [domain] will be regarded as the instructions of the virtual character at this position, thus starting the game.
[0075] This embodiment can be applied to fields such as cultural exhibitions, education and training, interactive exhibitions, conference offices, and personal productivity. During the process of completing the preset interaction activity participated by multiple users, in order to increase the interactive fun and experience of the preset interaction activity, the intelligent projection device can be used to project and / or collect images of the target area range involved in the activity, and the target objects within the target area range can be determined through image collection.
[0076] S120. Based on the interaction model and the first target object, project the corresponding guiding virtual images to each first target object.
[0077] Among them, the interaction model is a model trained based on the content involved in the preset interaction activity. This model can project different virtual images, light beams, and give different guiding prompts according to different nodes of the progress of the preset interaction activity. Among them, the guiding virtual image contains guiding information and / or explanatory information, which can be used to explain the image of the first target object or guide the user to carry out the next activity, and can also be used to interpret the activity rules and / or activity content of the preset interaction activity.
[0078] Specifically, after obtaining the first target object, in this embodiment, the interactive model can be used to further identify the first target object to determine the guiding virtual image corresponding to the first target object. Exemplarily, in this embodiment, the guiding virtual image can be projected at a specific position around the first target object to explain the first target object for the convenience of the user's understanding, such as Figure 2 as shown. Figure 2 FIG. Figure 2 is an exemplary schematic diagram of the projection of the guiding virtual image provided by the embodiment of the present disclosure, where the first target object 20 is a card, the guiding virtual image 21 is an explanation of the card, and the guiding virtual image 21 moves following the position of the first target object 20.
[0079] In another embodiment, in order to create the atmosphere of the user in the activity, this embodiment can also project the guiding virtual image for creating the atmosphere within the target area range, such as Figure 3 as shown. Figure 3 FIG. Figure 3 is another exemplary schematic diagram of the projection of the guiding virtual image provided by the embodiment of the present disclosure, including the target area range 300, and the arrow virtual image created by the guiding virtual images 310 and 320, as Figure 3 shown. The guiding virtual image in this embodiment may not include guiding information and explanatory information, and is only used to create the atmosphere.
[0080] S130. Within the target area range, obtain the response information of each user based on the guiding virtual image.
[0081] Among them, the response information may include but is not limited to the user's facial information, gesture information, action information, and voice information.
[0082] Specifically, after projecting the corresponding guiding virtual image onto the first target object, the user can obtain the necessary information for participating in the preset interactive activity according to the guiding virtual image and make corresponding responses according to the necessary information. For example, if the preset interactive activity is a Werewolf game in an offline board game scenario, the guiding information displayed on the guiding virtual image may be that the user needs to perform specific actions in a specific scenario according to the card content. Exemplarily, it can tell the user to open their eyes when they hear the host say "daytime".
[0083] S140. Based on the interactive model and the response information, determine the first guiding voice.
[0084] Among them, the first guiding voice is used to promote the preset interactive activity.
[0085] Since the interaction model is a model trained specifically according to preset interaction activities, any situation during the user participation process has been rehearsed, and it knows what operations to perform under what circumstances. Therefore, in this embodiment, the first guiding voice can be determined through the interaction model and the response information, and the next process can be carried out through the first guiding voice. For example, if the preset interaction activity is a Werewolf game in an offline board game scenario, the response information can be the voice information of one of the players, player A. After player A finishes speaking, this embodiment can give the first guiding voice as "Now please player B speak".
[0086] S150. Determine a first target user through the first guiding voice, and project a first light beam towards the first target user.
[0087] Among them, the first target user can be the user who needs to provide the light beam. For example, when player A is speaking, project the first light beam towards player A. The first light beam can be a light beam set according to the preset interaction activity and specifically adapted to this activity. It can be a single-color light beam or a light beam with changing colors, and its light beam range is not limited.
[0088] In this embodiment, during the process of completing the preset interaction activity participated by multiple users, the target objects within the target area are determined, and based on the interaction model and the first target objects, the corresponding guiding virtual images are projected towards each first target object; within the target area, the response information made by each user based on the guiding virtual image is obtained; based on the interaction model and the response information, the first guiding voice is determined; the first target user is determined through the first guiding voice, and a first light beam is projected towards the first target user. By constructing a three-dimensional interaction system that links vision, hearing, and behavior, the immersive participation experience in a multi-user scenario can be effectively improved, and the interest and collaboration of the interaction activity can be significantly enhanced.
[0089] In the embodiment of the present disclosure, after projecting the first light beam towards the first target user, it further includes: obtaining the first response information of the first target user in real time; when determining that the first response information is an end information, determining a second target user based on the interaction model; moving the first light beam to the second target user, and outputting a second guiding voice, where the content of the second guiding voice is different from that of the first guiding voice.
[0090] Among them, the first reaction information is the reaction information made by the first target user, including but not limited to the facial information, gesture information, action information and voice information of the first target user. The end information is to adjust the local or cloud AI large model Agent by setting specific prompt words and knowledge base when training the interaction model, so that the semantic model branch in the interaction model can accept more diverse language expressions, and finally classify it into several conclusions, and process it into executable instructions for the content program. For example, if the preset interactive activity is a werewolf killing game in an offline board game scene, the end information can be words such as "I'm done", "Pass!", "I'm done speaking", etc. Through the end information, this embodiment can analyze the instruction "the current player has finished speaking and needs to execute the next player's speech".
[0091] At the same time, this embodiment can also determine the second target user based on the interaction model, move the first light beam to the second target user, and output the second guiding voice to remind the second target user to speak. Specifically, this embodiment shines a spotlight in front of the player whose turn it is to speak, and gives a voice prompt "Player No. 4, please speak". After player No. 4 has explained his point of view, he finally says "I'm done", then moves the spotlight to the next player, and gives a voice prompt "Player No. 5, please speak".
[0092] This embodiment can provide roles that previously required real people to play, such as game hosts, DM (Dungeon Master), God, and game commentators, for offline table game scenes and multi-person party entertainment scenes. At the same time, it can create a more immersive game entertainment experience through sound and light and shadow, allowing every user to participate in and fully experience the entire game process.
[0093] In the above embodiment, step S110 can also be performed as follows: Figure 4 As shown, the specific steps include:
[0094] S111, acquiring an original image within the target area.
[0095] Wherein, the original image contains multiple targets.
[0096] Specifically, this embodiment can use any one or a combination of cameras to collect images of the target area as the original image. For example, if the preset interactive activity is a Werewolf game in an offline board game scene, this embodiment can capture RGB images of cards and three-dimensional objects (board game props) on the table as the original image, and by dedistorting the original image, sufficient details and feature points of the target object can be obtained.
[0097] S112: Recognize the original image through a recognition model to obtain two-dimensional image coordinates of the multiple targets.
[0098] Among them, the recognition model is used to recognize the multiple target objects. In the process of training the recognition model in this embodiment, any neural network model in the prior art can be adopted, and a large number of image samples containing specific objects are collected; the light and shadow variables and background environmental noises are added to the image samples as a sample set; the sample set is labeled and then input into the neural network model for training to adjust the model training parameters. Through repeated tests and verifications in this embodiment, the optimal inference and recognition results can be obtained, and a mature recognition model can be obtained. Exemplarily, this embodiment can adopt the YOLOv11 pre-training model for training.
[0099] In this embodiment, by training the recognition model, when collecting image samples, by adding light and shadow variables and background environmental noises, the training set is made more in line with the actual scenario. By improving the sample diversity, not only can the model training difficulty be effectively improved, but also the recognition accuracy of the model can be improved.
[0100] S113. Obtain the depth image of the target object within the target area range.
[0101] In this embodiment, a depth camera can be used to capture the depth image of the object within the target area range.
[0102] S114. Determine the true three-dimensional coordinates of the contour points of the multiple target objects through the depth image and the two-dimensional image coordinates of the multiple target objects.
[0103] Specifically, in this embodiment, a depth camera can be used to capture the depth map of the object within the target area range, convert the pixel points included in the depth map into the actual distance from the actual object to the camera lens plane, and then align the depth map with the original image captured in step S111, so as to obtain the distance from the contour vertices of each target object in the original image to the camera plane.
[0104] Specifically, in this embodiment, the true three-dimensional coordinates of the contour points of the multiple target objects can also be calculated based on the depth image obtained in the above steps and the two-dimensional image coordinates of each target object.
[0105] Exemplarily, if there is a pixel point A(u, v) in the depth image, assuming that its corresponding depth value is D (in meters), the camera internal parameter matrix is the following formula (1):
[0106]
[0107] Among them, f x and f y are the camera focal lengths, and c x and c y are the camera optical center coordinates.
[0108] Then the coordinates of the pixel point A in the real three-dimensional space as (X, Y, Z) can be calculated by the following formulas (2) - (4):
[0109]
[0110] Z = D (4)
[0111] In this embodiment, the camera can be used as the reference system origin. Through formula conversion, the three-dimensional coordinates of the contour vertices of each target object in the real space can be obtained.
[0112] S115. Determine each target object and its corresponding target label according to the real three-dimensional coordinates of each target object contour point.
[0113] Among them, the target label is used to determine the guiding virtual image to be projected by each target object by comparing with the interaction model. For example, the target label can be the type label of the target object, such as its name, size, rotation angle, and central three-dimensional coordinates, etc.
[0114] In this embodiment, after obtaining the real three-dimensional coordinates of each target object contour point, the contour points can be arranged in an orderly manner. By the orderly arranged contour points, it can be recognized what each target object is, and the target label of each target object can be determined. Then, in this embodiment, based on the comparison between the target label and the interaction model, the guiding virtual image to be projected by each target object is determined.
[0115] This embodiment can achieve high-precision recognition and positioning capabilities for two-dimensional and three-dimensional objects in the real space, providing a necessary guarantee for subsequent realization of immersive experiences and interaction pleasures for multiple users.
[0116] In the embodiment of the present disclosure, projecting the corresponding guiding virtual image to each first target object based on the interaction model and the first target object includes: restoring each target object in the virtual space based on the real three-dimensional coordinates of each target object contour point; determining the guiding virtual image to be projected by the first target object based on the target label and the interaction model; projecting the guiding virtual image in the surrounding area of the first target object. Among them, the guiding virtual image can be position-tracked and content-updated according to the first target object.
[0117] In this embodiment, in a virtual 3D simulation software (such as a game engine), the range of the target area and the characteristic information of each target can be restored 1:1 in the virtual space. Then, using the target label as an index, the guiding virtual image to be projected by the first target is determined within the interaction model, and the guiding virtual image is projected onto the surrounding area of the first target. Moreover, when the first target moves, the guiding virtual image can perform position tracking, follow the first target to move, and can also update the guiding virtual image according to actual needs.
[0118] For example, if the preset interaction activity is the Werewolf game in an offline board game scenario, then this embodiment can restore the desktop and the characteristic information of the recognition object 1:1, produce the virtual content required by the scenario, such as the card usage instruction panel, that is, the user interface (UI), and follow the recognition object at an appropriate angle and size. This embodiment can also, after calibrating the parameters of the virtual camera for zooming, offset, and tilt, project the guiding virtual image onto the desktop through a short-focus high-definition optical machine to obtain an interactive experience of the fusion of light and shadow between the virtual and the real. For example, before the Werewolf game officially starts, the user can place the card on the desktop, and project the usage instructions of the card around the card to let the players understand the game rules.
[0119] This embodiment can restore 2D and 3D objects in the real world to the virtual world in a 1:1 manner, and then project the virtual content onto the corresponding objects through software calculation to achieve a visually immersive experience of the fusion of the virtual and the real, improving the user's interactive pleasure and immersion experience.
[0120] In the embodiment of the present disclosure, when determining that the first response information is end information, determining a second target user based on the interaction model includes: when the first response information is the voice information of the first target user, determining the semantic information of the voice information through a semantic model; based on the semantic information, determining whether the first target user terminates speaking; and, based on the interaction model and the semantic information, determining the second target user.
[0121] Among them, the semantic model is a model for identifying the voice information of the user, which can be an independent model or a branch module in the interaction model.
[0122] Since the interaction model is a model trained based on the content involved in the preset interaction activity and can promote the progress of the activity in various situations, in this embodiment, when the first response information is the voice information of the first target user, the semantic information of the voice information can be determined through the semantic model, and then based on the semantic information, it can be determined whether the first target user terminates speaking. In the case of determining that the first target user terminates speaking, the second target user is determined based on the interaction model to promote the progress of the activity.
[0123] For example, when the preset interactive activity is a Werewolf game, in the scenario where players take turns speaking during the day in the game, this embodiment can determine the speech semantics of the current speaker. For example, the speech "I'm done speaking" will only take effect when spoken by the correct speaker, and then the right to speak will be passed to the next player. At the same time, this embodiment can also combine light and shadow interaction to move the light beam from the current speaker to the next player.
[0124] This embodiment can act as a host and promote the progress of the activity according to the user's speech information and interaction model, enabling each user to participate and enhancing the interactive experience.
[0125] In the embodiments of the present disclosure, it further includes: when the speech information of the third target user is collected, projecting a second light beam to the third target user.
[0126] Among them, the second light beam can be the same as the first light beam or different from the first light beam. Specifically, in this embodiment, when the first target user and / or the second target user is speaking, when it is detected that the third target user is speaking, that is, when it is detected that someone else is speaking, a second light beam can be projected to the third target user to attract the attention of others.
[0127] In the embodiments of the present disclosure, determining the first guiding speech based on the interaction model and the reaction information includes: determining the node advancement condition based on the interaction model; when the reaction information meets the node advancement condition, determining the first guiding speech.
[0128] Among them, the node advancement condition can be a condition for promoting the progress of the preset interactive activity.
[0129] Specifically, the node advancement condition of this embodiment can be a specific condition set according to different contents of the preset interactive activity. When the collected user reaction information meets the node advancement condition, the first guiding speech is determined and output to guide the user to carry out subsequent activities.
[0130] In the embodiments of the present disclosure, the reaction information includes facial information, gesture information, and / or speech information, and the reaction information meeting the node advancement condition includes at least one of the following: obtaining the facial information of each user, when there is at least one user's facial information that meets the facial advancement condition; obtaining the gesture information of each user, when there is at least one user's gesture information that meets the gesture advancement condition; obtaining the speech information of each user, when there is at least one user's speech information that meets the speech advancement condition.
[0131] Specifically, the reaction information of the user in this embodiment may include facial information, gesture information, and / or voice information. Exemplarily, for ease of understanding, this embodiment takes the Werewolf game as an example for illustration. For the facial recommendation condition, this embodiment collects the facial information of all users. It can be determined that the facial recommendation condition is met when there is one user opening or closing their eyes, or when several specific users open or close their eyes, or when all users close their eyes. This embodiment does not limit the specific facial recommendation condition.
[0132] Exemplarily, for the gesture promotion condition, this embodiment collects the gesture information of all users. It can be determined that the gesture recommendation condition is met when there is one user / multiple users raising their hands. This embodiment does not limit the specific gesture recommendation condition. For the voice promotion condition, this embodiment collects the voice information of all users. It can be determined that the voice recommendation condition is met when there is a user's voice information as the end information.
[0133] In this embodiment, by capturing and recognizing the states of two-dimensional and three-dimensional objects and combining the conditions required for the game scene to progress, it is determined whether to advance the game. For example, in the Werewolf game, everyone needs to close their eyes and turn the identity cards face up for the system to record their identities, and then they are required to turn the identity cards face down to avoid exposing their identities to other players. Only when all players meet the conditions required by the system will it advance to the next step.
[0134] This embodiment can not only interpret specific scenario games or content, but also combine real-world captured information for scene advancement, providing a more natural interaction experience through intelligent scene understanding.
[0135] Figure 5 It is a schematic structural diagram of an interactive projection device provided by an embodiment of the present disclosure. The device specifically includes:
[0136] A target object determination module 510, configured to determine a target object within a target area range during the process of completing a preset interaction activity participated by multiple users, where the target object includes a first target object, and the first target object corresponds to a user one by one;
[0137] An image projection module 520, configured to project the corresponding guiding virtual image to each first target object based on an interaction model and the first target object, where the interaction model is a model trained based on the content involved in the preset interaction activity;
[0138] An acquisition module 530, configured to acquire, within the target area range, the reaction information made by each user based on the guiding virtual image;
[0139] A voice determination module 540, configured to determine a first guiding voice based on an interaction model and the reaction information, where the first guiding voice is used to advance the preset interaction activity;
[0140] A light beam projection module 550, configured to determine a first target user through the first guiding voice, and project a first light beam to the first target user.
[0141] In an implementable manner, it further includes:
[0142] A light beam movement module, configured to, after projecting the first light beam to the first target user, obtain first reaction information of the first target user in real time; when determining that the first reaction information is an end information, determine a second target user based on the interaction model; move the first light beam to the second target user, and output a second guiding voice, where the content of the second guiding voice is different from that of the first guiding voice.
[0143] In an implementable manner, the target determination module 510 is specifically configured to: obtain an original image within the target area range, where the original image includes multiple target objects; identify the original image through an identification model to obtain two-dimensional image coordinates of the multiple target objects, where the identification model is used to identify the multiple target objects; obtain a depth image of the target objects within the target area range; determine real three-dimensional coordinates of contour points of the multiple target objects through the depth image and the two-dimensional image coordinates of the multiple target objects; and determine each target object and its corresponding target label according to the real three-dimensional coordinates of each target object contour point, where the target label is used to determine a guiding virtual image to be projected on each target object by comparing with the interaction model.
[0144] In an implementable manner, the image projection module 520 is specifically configured to: restore each target object in a virtual space based on the real three-dimensional coordinates of each target object contour point; determine a guiding virtual image to be projected on the first target object based on the target label and the interaction model; and project the guiding virtual image on a peripheral area of the first target object.
[0145] In an implementable manner, the light beam movement module is specifically configured to: when the first reaction information is voice information of the first target user, determine semantic information of the voice information through a semantic model; determine whether the first target user terminates speaking based on the semantic information; and determine the second target user based on the interaction model and the semantic information.
[0146] In an implementable manner, it further includes:
[0147] A projection module, configured to project a second light beam to the third target user when voice information of the third target user is collected.
[0148] In an implementable manner, the voice determination module 540 is specifically configured to: determine a node advancement condition based on an interaction model; and determine the first guiding voice when the response information meets the node advancement condition.
[0149] In an implementable manner, the response information meeting the node advancement condition includes at least one of the following: obtaining facial information of each user, when there is at least one user's facial information that meets the facial advancement condition; obtaining gesture information of each user, when there is at least one user's gesture information that meets the gesture advancement condition; obtaining voice information of each user, when there is at least one user's voice information that meets the voice advancement condition.
[0150] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium.
[0151] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0153] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a disk, optical disc, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0154] Computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 executes the various methods and processes described above, such as the interactive projection method. For example, in some embodiments, the interactive projection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the interactive projection method described above can be executed. Alternatively, in other embodiments, computing unit 601 can be configured to execute the sound effect correction method by any other suitable means (e.g., by means of firmware).
[0155] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0156] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0157] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0158] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0159] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0160] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0161] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0162] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one such feature. In the description of this disclosure, "a plurality" means two or more, unless otherwise specifically defined.
[0163] As described above, the above are only specific implementation manners of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.
Claims
1. An interactive projection method, characterized in that, The method includes: During the process of completing a preset interaction activity participated by multiple users, determining an object within a target area, where the object includes a first object, and the first object corresponds to each user one by one; Based on an interaction model and the first object, projecting a corresponding guiding virtual image onto each first object, where the interaction model is a model trained based on the content involved in the preset interaction activity; Within the target area, obtaining response information of each user based on the guiding virtual image; Based on the interaction model and the response information, determining a first guiding voice, where the first guiding voice is used to promote the preset interaction activity; Determining a first target user through the first guiding voice, and projecting a first light beam onto the first target user.
2. The method according to claim 1, characterized in that, After projecting the first light beam onto the first target user, it further includes: Obtaining the first response information of the first target user in real time; When determining that the first response information is an end message, determining a second target user based on the interaction model; Moving the first light beam to the second target user, and outputting a second guiding voice, where the content of the second guiding voice is different from that of the first guiding voice.
3. The method according to claim 2, wherein, The determining of the object within the target area includes: Obtaining an original image within the target area, where the original image contains multiple objects; Identifying the original image through an identification model to obtain two-dimensional image coordinates of the multiple objects, where the identification model is used to identify the multiple objects; Obtaining a depth image of the objects within the target area; Determining the true three-dimensional coordinates of the contour points of the multiple objects through the depth image and the two-dimensional image coordinates of the multiple objects; According to the true three-dimensional coordinates of the contour points of each object, determining each object and its corresponding target label, where the target label is used to determine the guiding virtual image to be projected onto each object by comparing with the interaction model.
4. The method according to claim 3, characterized in that, The projecting of the corresponding guiding virtual image onto each first object based on the interaction model and the first object includes: Restoring each object in a virtual space based on the true three-dimensional coordinates of the contour points of each object; Determining the guiding virtual image to be projected onto the first object based on the target label and the interaction model; Projecting the guiding virtual image onto the surrounding area of the first object.
5. The method according to claim 2, wherein The determining of the second target user based on the interaction model when determining that the first response information is an end message includes: When the first response information is the voice information of the first target user, determining the semantic information of the voice information through a semantic model; Based on the semantic information, determining whether the first target user terminates speaking; and Determining the second target user based on the interaction model and the semantic information.
6. The method according to claim 2, wherein It further includes: When collecting the voice information of a third target user, projecting a second light beam onto the third target user.
7. The method according to claim 1, wherein The determining of the first guiding voice based on the interaction model and the response information includes: Based on the interaction model, determine the node advancement condition; When the response information meets the node advancement condition, determine the first guiding voice.
8. The method according to claim 7, wherein The response information includes facial information, gesture information, and / or voice information. The response information meeting the node advancement condition includes at least one of the following: Obtain the facial information of each user. When there is at least one user's facial information that meets the facial advancement condition; Obtain the gesture information of each user. When there is at least one user's gesture information that meets the gesture advancement condition; Obtain the voice information of each user. When there is at least one user's voice information that meets the voice advancement condition.
9. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-8.