Method and device for controlling virtual object in application program, equipment and medium
By detecting the matching of user action sequences and predetermined action sequences, controlling virtual objects in virtual reality and augmented reality applications, the problem of single action recognition and boring control methods in the prior art is solved, and more flexible and convenient user interaction is achieved.
Patent Information
- Application Number
- CN202311688650.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-10
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art recognizes virtual objects in virtual reality and augmented reality applications with a single action and boring control method, lacking flexibility and convenience.
By detecting the matching of the user's action sequence and the predetermined action sequence, a virtual object is presented in the display area of the application, and the target virtual object is determined based on the user's interaction, and the corresponding operation is performed.
It realizes user interaction in a more flexible and effective way without interfering with the user experience, providing a richer way to control virtual objects.
Smart Images

Figure CN120122804A_ABST
Abstract
Description
Technical Field
[0001] Exemplary implementations of the present disclosure generally relate to application control, and particularly to methods, apparatuses, devices, and computer-readable storage media for controlling virtual objects in an application. Background Art
[0002] In recent years, technologies such as virtual reality (VR) and augmented reality (AR) have been widely studied and applied. By combining hardware devices and various technical means, they integrate virtual content and real scenes to provide users with unique sensory experiences. VR uses a computer to simulate a virtual world in three-dimensional space, providing users with immersive experiences in aspects such as vision, hearing, and touch. AR enables the real environment and virtual objects to be superimposed and coexist in the same space in real time. Currently, a variety of applications have been developed based on VR and AR, which can detect the actions of users of the application and then control various virtual objects in the application based on the detected actions. It is desirable to provide a more flexible and convenient way to control various virtual objects. Summary of the Invention
[0003] In a first aspect of the present disclosure, there is provided a method for controlling virtual objects in an application. The method includes: in response to determining that a first action sequence of a user of the application matches a first predetermined action sequence, presenting a set of virtual objects in a first format in a display area of the application; determining a target virtual object in the set of virtual objects based on user interaction of the user; and in response to determining that a second action sequence of the user matches a second predetermined action sequence, performing an operation corresponding to the target virtual object.
[0004] In a second aspect of the present disclosure, there is provided an apparatus for controlling virtual objects in an application. The apparatus includes: a first presentation module configured to present a set of virtual objects in a first format in a display area of the application in response to determining that a first action sequence of a user of the application matches a first predetermined action sequence; a determination module configured to determine a target virtual object in the set of virtual objects based on user interaction of the user; and an execution module configured to perform an operation corresponding to the target virtual object in response to determining that a second action sequence of the user matches a second predetermined action sequence.
[0005] In a third aspect of the present disclosure, there is provided an electronic device. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor causes the processor to implement the method according to the first aspect of the present disclosure.
[0007] It should be understood that the content described in this section is not intended to limit the key features or important features of the implementation manners of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the following, with reference to the accompanying drawings and the following detailed description, the above and other features, advantages and aspects of the various implementation manners of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 A block diagram showing an application environment according to some exemplary implementation manners of the present disclosure is shown;
[0010] Figure 2 A block diagram showing a device for controlling a virtual object in an application according to some implementation manners of the present disclosure is shown;
[0011] Figure 3 A block diagram showing a device for selecting a virtual object according to some implementation manners of the present disclosure is shown;
[0012] Figure 4 A block diagram showing the structure of a wearable device according to some implementation manners of the present disclosure is shown;
[0013] Figure 5 A block diagram showing performing an operation by using a selected virtual object according to some implementation manners of the present disclosure is shown;
[0014] Figure 6A A block diagram showing a device for selecting a predetermined action sequence of an operation according to some implementation manners of the present disclosure is shown;
[0015] Figure 6B A block diagram showing a device for releasing a predetermined action sequence of an operation according to some implementation manners of the present disclosure is shown;
[0016] Figure 7 A block diagram showing a device for specifying a predetermined action sequence according to some implementation manners of the present disclosure is shown;
[0017] Figure 8 A flowchart showing a method for controlling a virtual object in an application according to some implementation manners of the present disclosure is shown;
[0018] Figure 9A block diagram of an apparatus for controlling virtual objects in an application according to some implementations of the present disclosure is shown; and
[0019] Figure 10 A block diagram of a device capable of implementing multiple implementations of the present disclosure is shown. Detailed Implementations
[0020] Implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. On the contrary, these implementations are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the implementations of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". There may also be other explicit and implicit definitions hereinafter. As used herein, the term "model" may represent the association relationship between various data. For example, the above association relationship can be obtained based on various technical solutions known currently and / or to be developed in the future.
[0022] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0023] It can be understood that before using the technical solutions disclosed in the various implementations of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.
[0024] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0025] As an optional but non-limiting implementation, in response to receiving an active request from a user, a way to send a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It can be understood that the above notification and user authorization acquisition process is only illustrative and does not limit the implementation of the present disclosure. Other ways that comply with relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0027] The term "in response to" used herein represents a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the execution timing of the subsequent actions executed in response to the event or condition and the time when the event occurs or the condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent actions can be executed immediately when the event occurs or the condition is established; while in other cases, the subsequent actions can be executed after a period of time after the event occurs or the condition is established.
[0028] Example environment
[0029] Figure 1 FIG. shows a block diagram of an application environment 100 according to an exemplary implementation of the present disclosure. The application environment 100 includes a user 130 and his / her wearable device 110. In this application environment 100, the wearable device 110 installs and runs an application program that supports a virtual environment, thereby supporting the display area 120 of the application program to be presented at the wearable device 110 or by the wearable device 110 to the user 130.
[0030] In some implementations, the display area 120 can provide a virtual environment (e.g., a VR / AR environment). As Figure 1 shown, when the user 130 wears the wearable device 110, the display area 120 can provide a representation including one or more virtual objects and / or real objects in the real world. The virtual object can include a virtual representation corresponding to a real object in the real world and / or an object only applicable to the virtual environment. Additionally, the user 130 can input instructions to the application program to control the operation of the application program.
[0031] Currently, a variety of application programs have been developed based on VR / AR, which can detect the actions of the users of the application programs and then control various virtual objects in the application programs based on the detected actions. However, currently, only some simple actions can be recognized, and the control methods are single and boring. At this time, it is desired to provide a more flexible and convenient way to control various virtual objects.
[0032] Overview of controlling virtual objects
[0033] To at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for controlling virtual objects in an application is proposed. Generally speaking, actions performed by a user can be detected and tracking can be performed for various types of user interactions. If it is determined that the detected action meets a predetermined condition, a virtual object can be presented. Further, if it is found that the user's interaction meets a predetermined condition, the presented virtual object can be selected, and then an operation corresponding to the virtual object can be executed. In this way, the user does not have to hold a dedicated input device, but can provide user interaction in a more flexible and efficient manner without interfering with the user's experience of using the application.
[0034] It should be understood that the applications herein can include various types. For example, dedicated applications for implementing certain specific functions, such as a brush application for editing images, a text application for editing text, and so on. Alternatively and / or additionally, the application can further include system-level applications, such as an operating system, and so on.
[0035] For ease of description, the detailed process of controlling virtual objects will be described by taking a brush application as an example of the application. Refer to Figure 2 Describe the outline according to an exemplary implementation of the present disclosure, which Figure 2 shows a block diagram 200 for controlling virtual objects in an application according to some implementations of the present disclosure. As Figure 2 shown, the display area 120 can be the interface of a brush application, and the user can interact with the display area 120 to select various brushes and / or colors, and then use the selected tools to perform painting operations.
[0036] As Figure 2 shown, an action sequence of the user of the application can be detected (for ease of description, this action sequence can be referred to as the first action sequence). It can be determined whether the detected action sequence matches a predetermined action sequence (for example, the predetermined action sequence for presenting the menu of virtual objects can be referred to as the first predetermined action sequence 210). Here, the first predetermined action sequence 210 can be an action sequence default in the application or an action sequence customized by the user. For ease of description, the first predetermined action sequence 210 can be an action sequence of the user's hand and includes: [Place the hand horizontally, palm up, and spread the palm].
[0037] It should be understood that multiple actions in the action sequence herein can respectively correspond to multiple time points, and there is a chronological relationship between the multiple time points. For example, placing the hand horizontally can correspond to time point t1, turning the palm upward can correspond to time point t2, and spreading the palm can correspond to time point t3. At this time, t1 is before t2, and t2 is before t3.
[0038] It should be understood that although the action sequence in the above text includes multiple gestures, alternatively and / or additionally, the action sequence can include only one gesture. In the context of the present disclosure, a predetermined action sequence can include one or more gestures, where the gesture can represent the posture of the user's hand. For example, the hand placement horizontally, turning the palm upward, spreading the palm, etc. described above.
[0039] Alternatively and / or additionally, the gesture can include the position of the hand, that is, the action can be described solely by using the position of the hand. For example, when the user's hand is at different positions, it can represent different actions. At this time, the user can move the hand along a spatial trajectory to form an action sequence. Alternatively and / or additionally, the action can further involve the gesture and position of the hand, that is, both the gesture of the hand and the position of the hand are used to describe the gesture. For example, when the user turns the palm upward and the hand is in front of the chest, it can represent one action; when the user turns the palm upward and the hand is at the side of the body, it can represent another action. In this way, the action can be described in a more precise manner by using the action and position of the hand.
[0040] It should be understood that although the above text describes the action sequence with hand actions as specific examples, alternatively and / or additionally, at least one action in the action sequence can include a body action performed by at least one body part of the user. For example, a body action that can be performed by the user's fingers, hand, upper arm, head, leg, torso, and / or other one or more body parts or a combination thereof. Here, the body action of the user can be detected separately. For example, the action can be represented solely by using the posture of the body part, solely by using the position of the body part, or by using a combination of the posture and position of the body part. In this way, the actions of multiple body parts of the user can be comprehensively collected.
[0041] According to an example implementation of the present disclosure, the action may include an interaction action of the user with an accessory device of the wearable device. For example, the user may operate buttons, rollers, etc. in an accessory device such as a handle to trigger the interaction action. In one example, the user's interaction action may be detected separately. During the user's operation of the handle, the user may move the position of the handle or keep the position of the handle unchanged. At this time, only the interaction action of the user with the buttons, rollers, etc. may be detected, and the change in the position of the user's hand may be ignored. In another example, both the user's interaction action and body movement may be detected. For example, the user may hold the handle and move the handle along a certain trajectory, and when the handle reaches a predetermined position, press the button. At this time, the action sequence may include a combination of body movement and interaction action. In this way, a richer action sequence may be defined, thereby avoiding the situation of misjudgment that may occur in a simple action sequence.
[0042] If it is determined that the detected first action sequence of the user matches the first predetermined action sequence 210, a menu may be presented in the display area 120 of the application program, that is, a set of virtual objects 220 may be presented in the first format. As Figure 2 shown, in the scenario of a painting application, a set of virtual objects 220 may be presented in black in a circular menu, and at this time, a set of virtual objects 220 may include various different types of paintbrushes, for example, a pencil, a pen, an oil paintbrush, and so on. Alternatively and / or additionally, a set of virtual objects 220 may relate to paintbrushes of different colors, for example, red, yellow, blue, and so on. It should be understood that although the set of virtual objects shown above includes multiple virtual objects, alternatively and / or additionally, a set of virtual objects may include only one virtual object. In other words, a set of virtual objects may include at least one virtual object.
[0043] Furthermore, it may be determined whether the user interaction of the user matches a certain virtual object in a set of virtual objects. It should be understood that the user interaction here may be implemented based on various methods. For example, the user interaction may include user interaction based on at least any one of the following: the user's line of sight direction, the user's eye movement, the user's voice interaction, and the user's body interaction, and so on. For the convenience of description, only the user's line of sight direction will be taken as an example of the user interaction hereinafter. Alternatively and / or additionally, the user may use a blinking action, say "select" or "page turning", use the other hand to perform a selection action, etc. to implement the user interaction.
[0044] According to an example implementation of the present disclosure, when using the line-of-sight direction to represent user interaction, it is possible to determine whether the user is gazing at a certain virtual object in the menu, and then determine the virtual object and perform an operation corresponding to the virtual object. If it is found that the line-of-sight direction 240 matches the target virtual object 230 in a group of virtual objects 220, and the target virtual object 230 corresponds to a red paintbrush, then the drawing operation can be performed using the red paintbrush.
[0045] Alternatively and / or additionally, to facilitate user operation, the target virtual object 230 can be highlighted in the display area 120. For example, the target virtual object 230 can be presented in a second format different from the first format (e.g., in a different color, using a bounding box selection, changing the display size, changing the display position, or adding other display effects, etc.).
[0046] See Figure 3 for more details, the Figure 3 block diagram 300 for selecting a virtual object according to some implementations of the present disclosure is shown. As Figure 3 shown, it is possible to further detect other action sequences of the user after the first action sequence (e.g., simply referred to as the second action sequence). Subsequently, it can be determined whether the second action sequence of the user matches another action sequence (e.g., a second predetermined action sequence for selecting a virtual object can be called). In the context of the present disclosure, the second predetermined action sequence can be an action sequence of the user's hand and includes: [hand placed horizontally, palm down, fist clenched]. If it is determined that the second action sequence matches the second predetermined action sequence 310, then the target virtual object 230 can be selected. As Figure 3 shown by the dashed line in, at this time, the group of virtual objects 220 in the circular menu and the already selected target virtual object 230 can no longer be presented.
[0047] According to an example implementation of the present disclosure, the end action in the first predetermined action sequence can be the same as the start action in the second predetermined action sequence. At this time, the second predetermined action sequence can include: [palm spread, palm down, fist clenched]. At this time, there are no other actions between the first predetermined action sequence and the second predetermined action sequence. In this way, the user's actions can be detected in a more continuous manner, thereby avoiding an abnormal state of accidentally exiting the virtual object caused by some irrelevant actions of the user.
[0048] In this way, user interaction can be provided in a more flexible and effective manner without disturbing the user's experience of using the application. At this time, the user does not have to hold any input device, but can implement user interaction only based on tracking the user's body movements and eye movements. At this time, the user's eye movements and body movements can cooperate with each other, that is, clench the fist when the user's eyes are gazing at the target virtual object, thereby reducing the risk of misidentifying the user's intention.
[0049] Detailed process of controlling virtual objects
[0050] The overview of an example implementation according to the present disclosure has been described. In the following, more details of controlling virtual objects will be described. According to an example implementation of the present disclosure, when the user has selected a target virtual object, the selected target virtual object can be presented in the display area in a third format different from the first format and the second format. As Figure 3 shown, a prompt 320 can be provided in the display area 120 to inform the user that a certain virtual object has been selected. Alternatively and / or additionally, the selected virtual object can be presented at a position near the user's hand. For example, the selected paintbrush can be presented at the position of the user's finger, and so on. In this way, the user can be informed of the currently selected virtual object, thereby reminding the user that the virtual object can be used to perform subsequent actions.
[0051] According to an example implementation of the present disclosure, the above process can be completed by the wearable device 110. Figure 4 FIG. 400 is a block diagram showing the structure of a wearable device according to some implementations of the present disclosure. As Figure 4 shown, the wearable device 110 can include a computing device 410, an environmental sensor device 420, and an eye tracking sensor 430.
[0052] Here, the computing device 410 can include various computing devices capable of providing computing capabilities, and an application 412 can run on the computing device, and then provide a display interface of the application in the display area 120. The environmental sensor 420 in the computing device 410 can be deployed outside the wearable device 110 and collect environmental data around the user. For example, image data and / or other data within a 360-degree range around the user can be collected to identify the action sequence 422 performed by the user. The eye tracking sensor 430 in the computing device 410 can be deployed along the direction towards the user's eyes to collect data and then use the collected data to detect and track the user's line of sight direction 432 and / or other data.
[0053] According to an example implementation of the present disclosure, the processes described above can be performed in a single wearable device 110. In this way, the user can wear the wearable device 110 and move around at will, thereby providing a richer user experience. For example, the paintbrush application herein can be an application that supports the user to paint in a three-dimensional space. At this time, the user can walk to the position where they expect to paint in the physical space and complete the painting operation. For example, the user can walk to a wall position in the room and paint a landscape painting on the wall. Or, for example, the user can walk to another wall position in the room and paint a portrait on the wall.
[0054] Alternatively and / or additionally, the computing device 410, the environmental sensor 420, and the eye tracking sensor 430 can be located at different positions. As long as they can communicate with the computing device 410, the environmental sensor 420, and the eye tracking sensor 430 via a network. It should be understood that although Figure 4 only a wearable device is schematically shown, alternatively and / or additionally, the wearable device can have other associated accessory devices. For example, a handle and / or other accessory devices can be used to cooperate with the wearable device to achieve more functions.
[0055] According to an example implementation of the present disclosure, a first action sequence performed by the user can be presented in the display area 120. The respective action sequences can be presented in a variety of ways. For example, the action sequence of the user can be displayed based on the perspective function of the head-mounted device, or the image of a virtual object that can represent the action sequence can be simulated using the relevant data of the action sequence to present the first action sequence, and so on. Further, a set of virtual objects can be presented near the first action sequence of the user. Return Figure 2 , a set of virtual objects 220 can be presented above the user's palm. Here, the set of virtual objects 220 can be presented in a translucent manner and / or an opaque manner.
[0056] At this time, the circular menu representing multiple paintbrushes can move accordingly as the user's hand moves. In other words, the presentation position of the set of virtual objects 220 in the display area 120 corresponds to the presentation position of the first action sequence in the display area 120. For example, when the user's hand moves to the right, the circular menu will move to the right as the user's hand moves, and when the user's hand moves to the left, the circular menu will move to the left as the user's hand moves, and their relative positions remain unchanged. In this way, there can be more interaction between the virtual objects in the application program and the user's actions, thus providing a more realistic visual effect.
[0057] According to an example implementation of the present disclosure, in the process of determining whether the line of sight direction matches the target virtual object, the reference line of sight direction of the user can be determined based on the eye position of the user and the virtual object position of the target virtual object. During the running of the application, multiple coordinate systems may be involved. For example, the world coordinate system can be used to describe the physical world where the user is located, the virtual coordinate system can be used to describe the three-dimensional virtual world corresponding to the physical world, and the screen coordinate system can be used to describe the two-dimensional display area 120 corresponding to the three-dimensional virtual world (that is, the three-dimensional virtual world is projected onto the two-dimensional display area 120).
[0058] According to an example implementation of the present disclosure, the eye position of the user and the virtual object position of the target virtual object can be determined based on various mapping relationships that are currently known and / or will be developed in the future (for example, using a transformation matrix), and then the reference line of sight direction can be determined. The reference line of sight direction here can represent the line of sight direction that should be followed when the user gazes at the target virtual object. Further, the line of sight direction of the user collected by the eye tracking sensor can be compared with the reference line of sight direction. If it is determined that the line of sight direction matches the reference line of sight direction, it can be determined that the line of sight direction matches the target virtual object, that is, at this time the user is gazing at the target virtual object.
[0059] According to an example implementation of the present disclosure, during the process of the user selecting a virtual object, the user can change the line of sight direction and look at other virtual objects. For example, in Figure 2 the user can look at another virtual object below the target virtual object 230. At this time, the other virtual object will be displayed in the second format, and the target virtual object 230 will resume the first format. In this way, the user can be allowed to flexibly select the virtual object of interest through the line of sight direction. This allows, without having to use hand and / or other limb movements, using eye movements to conveniently switch the virtual object to be selected, thereby improving the interaction efficiency.
[0060] The user can continuously adjust the viewing point and then perform actions to select the desired target virtual object. After the user has selected the target virtual object, subsequent action sequences can be further detected (for example, the subsequent action sequence after the second action sequence can be called the third action sequence). The environmental sensor 420 described above can be used to collect data, and then the collected data can be used to determine the third action sequence, and then an operation corresponding to the target virtual object can be performed based on the third action sequence.
[0061] See Figure 5 for more details. The Figure 5 shows a block diagram 500 of performing operations using the selected virtual object according to some implementations of the present disclosure. As Figure 5As shown, assume that the target virtual object selected by the user is the paintbrush 520. Then, in the display area 120, an image of the user's hand holding the paintbrush 520 can be presented. The user's hand can perform a third action sequence 510 (e.g., draw a curve) in the real world. At this time, a painting trajectory 530 will be presented in the display area 120, and the painting trajectory 530 is drawn using the paintbrush 520 selected by the user. Using the example implementation of the present disclosure, it is convenient for the user to implement corresponding functions using the selected virtual object. In this way, a richer interaction experience can be provided.
[0062] It should be understood that although the process of controlling a virtual object is described above using a paintbrush application as an example, alternatively and / or additionally, the application program can perform other functions. According to an example implementation of the present disclosure, a set of virtual objects can be a set of controls in the application program. In different application programs, the virtual objects can have different functions. For example, in a game program, a set of virtual objects can be different props in the game, such as stones, bows and arrows, etc. For another example, in a virtual assembly program, a set of virtual objects can be various repair tools, such as wrenches, hammers, etc. In this way, it is convenient to select the desired virtual object in a way that is more in line with the user's operating habits in the real world, thereby improving the interaction efficiency with the application program.
[0063] Alternatively and / or additionally, the virtual object can be presented in other graphical representations in the application program. For example, the virtual object can include presenting various virtual characters, virtual animals, and / or other virtual items in the application program. Using the example implementation of the present disclosure, action sequences can be used to operate multiple virtual objects in a more flexible manner, thus providing a more flexible operation method.
[0064] According to an example implementation of the present disclosure, different visual effects can be presented in the display area 120 based on the function of the virtual object. For example, in a game application, the user can pick up a stone and make a throwing action. At this time, the environmental sensor 420 can collect data, and further use the collected data to determine the user's action (e.g., projection angle, projection speed, etc.), and then generate and present the corresponding throwing trajectory of the stone. After the user has thrown the stone, the next stone can be automatically picked up, and at this time the user can perform the throwing action again. For another example, after the user has thrown the stone, the user needs to open the menu again to pick up the stone again. For another example, a predetermined action sequence of the hand and / or eye for picking up the next stone can be predefined, and so on. In a virtual assembly application, the user can use the selected tool to perform corresponding operations, such as using a wrench to rotate a screw, etc.
[0065] According to an example implementation of the present disclosure, a user may be allowed to customize a predetermined action sequence for performing various operations. Figure 6A FIG. 600A is a block diagram showing a predetermined action sequence for selecting an operation according to some implementations of the present disclosure. As Figure 6A shown, a predetermined action sequence for performing a selection operation 610 may be defined: [Place the hand horizontally, palm down, and clench the fist]. At this time, when it is detected that the user has performed the above action sequence, the target virtual object that the user is currently gazing at can be selected.
[0066] According to an example implementation of the present disclosure, a user is allowed to release the selected target object. For example, in the case where the user has selected a target virtual object, the user may wish to select another virtual object, or the user has finished using and wishes to replace it with another virtual object. At this time, a subsequent action sequence performed by the user can be detected (for example, this subsequent action sequence may be referred to as the fourth action sequence). It should be understood that the fourth action sequence may directly follow the second action sequence, which corresponds to the case where the user wishes to replace the virtual object when not using the target virtual object. For another example, the fourth action sequence may follow the third action sequence, corresponding to the case where the user wishes to replace the virtual object after having the target virtual object.
[0067] According to an example implementation of the present disclosure, it can be determined whether the user's fourth action sequence matches a predefined action sequence for releasing a virtual object (for example, referred to as the third predetermined action sequence). If it is determined that the user's fourth action sequence matches the third predetermined action sequence, the operation corresponding to the target virtual object can be exited. Refer to Figure 6B for more details. The Figure 6B FIG. 600B is a block diagram showing a predetermined action sequence for a release operation according to some implementations of the present disclosure. As Figure 6B shown, a predetermined action sequence for performing a release operation 620 may be defined: [Place the hand horizontally, palm up, and clench the fist]. At this time, when it is detected that the user has performed the above action sequence, the operation corresponding to the target virtual object can be exited, that is, the operation of drawing with a paintbrush can be exited. Using the example implementation of the present disclosure, user interaction can be performed in a more flexible and effective manner, thereby providing a richer user experience.
[0068] According to an example implementation of the present disclosure, although the above describes a case where the third predetermined action sequence is different from the first predetermined action sequence, alternatively and / or additionally, the third predetermined action sequence may be the same as the first predetermined action sequence. That is, the third predetermined action sequence may include the first predetermined action sequence. In this way, the action sequence for triggering the presentation of a set of virtual objects is the same as the action sequence for exiting the operation corresponding to the virtual object, and in this way, it can facilitate user memory.
[0069] According to an example implementation of the present disclosure, each of the above-described predetermined action sequences (for example, any one of the first predetermined action sequence, the second predetermined action sequence, and the third predetermined action sequence) may be a default action sequence set by the application program. Alternatively and / or additionally, the above-mentioned predetermined action sequence may be specified by the user. See Figure 7 for more details, the Figure 7 shows a block diagram 700 for specifying a predetermined action sequence according to some implementations of the present disclosure.
[0070] As Figure 7 shown, buttons 710, 720, and 730 for specifying each action sequence may be presented to the user in the display area 120. Specifically, a user interaction 740 for specifying a predetermined action sequence may be detected. In the case where the user interaction 740 is detected, a video for specifying a predetermined action sequence may be received. For example, the user may click the button 710 to record and / or load from the local or network a video including hand movements to specify the action sequence for opening the menu. Similarly, the user may click the button 720 to provide the action sequence for selecting a target virtual object from the menu, and the user may click the button 730 to provide the action sequence for releasing the selected target virtual object. Further, the predetermined action sequence may be determined from the video based on various methods known currently and / or to be developed in the future.
[0071] Using the example implementation of the present disclosure, the user is allowed to specify an action sequence with which they are familiar, thereby facilitating the improvement of interaction efficiency. With the improvement of the accuracy of various sensors, user actions can be recognized in a more accurate manner, thereby allowing the user to define an action sequence including fewer actions. For example, the user may specify: opening the menu with a "snapping fingers" action, selecting the target virtual object being gazed at with an "OK" action, and canceling the selected target virtual object with a "waving hand" action, and so on. In this way, a more convenient and friendly user experience can be provided.
[0072] According to an example implementation of the present disclosure, after presenting a menu of virtual objects, if it is found that the user has not responded for a long time, the menu may no longer be displayed. Specifically, if no action sequence matching the second predetermined action sequence is detected within a predetermined time period and the user's line of sight direction does not match any of the virtual objects in a group of virtual objects, the group of virtual objects may be removed from the display area. At this time, the user is not looking at the menu area, which indicates that the user may not wish to use the menu, so the menu can be removed.
[0073] If the user wishes to reopen the menu, the first predetermined action sequence described above may be re-executed to represent the menu again. If the user does not wish to present the menu, at this time the user has not selected any target virtual object, and the user may perform any subsequent action. In this way, the menu can be closed in time when the user does not need it, thereby preventing the error of selecting a certain virtual object due to the user's accidental operation.
[0074] Although the process for controlling virtual objects is described above by taking the actions of the user's hand as an example, in the context of the present disclosure, action sequences performed by one or more human body parts may be detected, and then the corresponding virtual objects may be controlled. At this time, each predetermined action sequence (for example, any one of the first predetermined action sequence, the second predetermined action sequence, and the third predetermined action sequence) includes an action sequence performed by any one of the following: an arm, both hands, and a single hand.
[0075] For example, the actions of the user's hand and / or the arm including the hand may be detected. At this time, each action involves more human body parts, so relatively complex predetermined action sequences may be predefined. In this way, richer interaction methods can be provided and the accuracy of determining user interaction can be improved. For another example, the actions of the user's both hands may be detected. Specifically, the user may use the left hand to perform the action of selecting a virtual object and use the right hand to perform the action of releasing the virtual object. In this way, the constraints on the user's actions can be reduced, and the user can select and / or release virtual objects according to their own habits.
[0076] For another example, the actions of the user's single hand may be detected. In this way, no other human body parts of the user are required to participate in the interaction, and the user's experience of using the application can be disturbed as little as possible. For example, the user's other hand may control other virtual objects in the application, and so on.
[0077] Using the example implementation of the present disclosure, user interaction can be provided in a more flexible and effective manner without disturbing the user's experience of using the application. At this time, the user does not have to hold any input device, but can implement user interaction only based on tracking the user's limb movements and eye movements. In this way, the user's limb movements and eye movements can cooperate with each other, thereby reducing the risk of misidentifying the user's intention.
[0078] Example process
[0079] Figure 8 FIG. 800 is a flowchart of a method for controlling virtual objects in an application according to some implementations of the present disclosure. At block 810, in response to determining that a first action sequence of a user of the application matches a first predetermined action sequence, a set of virtual objects is presented in a first format in a display area of the application. At block 820, a target virtual object in the set of virtual objects is determined based on the user's interaction. At block 830, in response to determining that a second action sequence of the user matches a second predetermined action sequence, an operation corresponding to the target virtual object is executed.
[0080] According to an example implementation of the present disclosure, the method further includes: presenting the first action sequence in the display area, wherein the presentation position of the set of virtual objects in the display area corresponds to the presentation position of the first action sequence in the display area.
[0081] According to an example implementation of the present disclosure, the method further includes: presenting the target virtual object in a second format different from the first format in the display area.
[0082] According to an example implementation of the present disclosure, the method further includes: presenting the selected target virtual object in a third format different from the first format and the second format in the display area.
[0083] According to an example implementation of the present disclosure, the method further includes: in response to detecting a third action sequence of the user, performing an operation corresponding to the target virtual object based on the third action sequence.
[0084] According to an example implementation of the present disclosure, the method further includes: in response to determining that a fourth action sequence of the user matches a third predetermined action sequence, exiting the operation corresponding to the target virtual object.
[0085] According to an example implementation of the present disclosure, the third predetermined action sequence includes the first predetermined action sequence.
[0086] According to an example implementation of the present disclosure, the method further includes at least any one of the following: presenting another target virtual object in a display area in a second format in response to determining that a user interaction matches another target virtual object in a group of virtual objects; and removing the group of virtual objects from the display area in response to not detecting an action sequence that matches a second predetermined action sequence within a predetermined period of time and the user interaction not matching any virtual object in the group of virtual objects.
[0087] According to an example implementation of the present disclosure, the predetermined action sequences in the first predetermined action sequence and the second predetermined action sequence are specified by a user.
[0088] According to an example implementation of the present disclosure, the method further includes: receiving a video for specifying a predetermined action sequence in response to detecting a user interaction for specifying a predetermined action sequence; and determining the predetermined action sequence from the video.
[0089] According to an example implementation of the present disclosure, any one of the first predetermined action sequence and the second predetermined action sequence includes an action sequence performed by any one of the following: an arm, both hands, and a single hand, and the ending action in the first predetermined action sequence is the same as the starting action in the second predetermined action sequence.
[0090] According to an example implementation of the present disclosure, any one of the first action sequence and the second action sequence is determined based on data collected from at least any one of the following: an environmental sensor in a wearable device worn by a user and an accessory device associated with the wearable device; the line-of-sight direction of the user interaction is determined based on data collected by an eye-tracking sensor in the wearable device; and the application runs on a computing device in the wearable device.
[0091] According to an example implementation of the present disclosure, the action sequence includes at least one action, and the action includes at least any one of the following: an interaction action of the user with an accessory device; and a body action of at least one body part of the user, and the body action is represented by at least any one of the following: the posture and position of at least one body part of the user.
[0092] According to an example implementation of the present disclosure, a group of virtual objects is a group of controls in an application, and the user interaction therein includes a user interaction based on at least any one of the following: the line-of-sight direction of the user, the eye movement of the user, the voice interaction of the user, and the limb interaction of the user.
[0093] Example devices and equipment
[0094] Figure 9A block diagram of an apparatus 900 for controlling virtual objects in an application according to some implementations of the present disclosure is shown. The apparatus 900 includes: a first rendering module 910 configured to render a set of virtual objects in a first format in a display area of the application in response to determining that a first action sequence of a user of the application matches a first predetermined action sequence; a determination module 920 configured to determine a target virtual object among the set of virtual objects based on user interaction of the user; and an execution module 930 configured to perform an operation corresponding to the target virtual object in response to determining that a second action sequence of the user matches a second predetermined action sequence.
[0095] According to an example implementation of the present disclosure, the apparatus further includes: a second rendering module configured to render the first action sequence in the display area, wherein a rendering position of the set of virtual objects in the display area corresponds to a rendering position of the first action sequence in the display area.
[0096] According to an example implementation of the present disclosure, the apparatus further includes: rendering the target virtual object in a second format different from the first format in the display area.
[0097] According to an example implementation of the present disclosure, the apparatus further includes: a third rendering module configured to render the selected target virtual object in a third format different from the first format and the second format in the display area.
[0098] According to an example implementation of the present disclosure, the apparatus further includes: an operation execution module configured to perform an operation corresponding to the target virtual object based on a third action sequence in response to detecting a third action sequence of the user.
[0099] According to an example implementation of the present disclosure, the apparatus further includes: an exit module configured to exit the operation corresponding to the target virtual object in response to determining that a fourth action sequence of the user matches a third predetermined action sequence.
[0100] According to an example implementation of the present disclosure, the third predetermined action sequence includes the first predetermined action sequence.
[0101] According to an example implementation of the present disclosure, the apparatus further includes: a fourth rendering module configured to render another target virtual object in the second format in the display area in response to determining that the user interaction matches another target virtual object among the set of virtual objects; and a removal module configured to remove the set of virtual objects from the display area in response to no action sequence matching the second predetermined action sequence being detected within a predetermined time period and the user interaction not matching any virtual object among the set of virtual objects.
[0102] According to an example implementation of the present disclosure, the predetermined action sequences in the first predetermined action sequence and the second predetermined action sequence are specified by the user.
[0103] According to an example implementation of the present disclosure, the apparatus further includes: a receiving module configured to receive a video for specifying a predetermined action sequence in response to detecting a user interaction for specifying a predetermined action sequence; and an action sequence determination module configured to determine a predetermined action sequence from the video.
[0104] According to an example implementation of the present disclosure, any one of the first predetermined action sequence and the second predetermined action sequence includes an action sequence performed by any one of the following: an arm, both hands, and a single hand, and the ending action in the first predetermined action sequence is the same as the starting action in the second predetermined action sequence.
[0105] According to an example implementation of the present disclosure, any one of the first action sequence and the second action sequence is determined based on data collected by at least any one of the following: an environmental sensor in a wearable device worn by the user, and an accessory device associated with the wearable device; the line-of-sight direction of the user interaction is determined by data collected by an eye-tracking sensor in the wearable device; and an application runs on a computing device in the wearable device.
[0106] According to an example implementation of the present disclosure, the action sequence includes at least one action, and the action includes at least any one of the following: an interaction action of the user with an accessory device; and a body action of at least one body part of the user, and the body action is represented by at least any one of the following: the posture and position of at least one body part of the user.
[0107] According to an example implementation of the present disclosure, a set of virtual objects is a set of controls in an application, and the user interaction therein includes user interaction based on at least any one of the following: the line-of-sight direction of the user, the eye movement of the user, the voice interaction of the user, and the limb interaction of the user.
[0108] Figure 10 The block diagram of a device 1000 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 10 The computing device 1000 shown is merely exemplary and should not constitute any limitation to the functions and scopes of the implementations described herein. Figure 10 The computing device 1000 shown can be used to implement the methods described above.
[0109] As Figure 10As shown, the computing device 1000 is in the form of a general-purpose computing device. The components of the computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, a memory 1020, a storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 may be an actual or virtual processor and is capable of performing various processes according to programs stored in the memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 1000.
[0110] The computing device 1000 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1020 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1030 may be removable or non-removable media and may include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data (such as training data for training) and can be accessed within the computing device 1000.
[0111] The computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 10 a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0112] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 1000 may be implemented in a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, the computing device 1000 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0113] The input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 1060 can be one or more output devices, such as a display, speaker, printer, etc. The computing device 1000 can also communicate with one or more external devices (not shown) as needed through the communication unit 1040. The external devices such as storage devices, display devices, etc., communicate with one or more devices that enable the user to interact with the computing device 1000, or communicate with any device that enables the computing device 1000 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0114] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0115] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0116] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0117] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0118] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0119] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of the technology without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the field of the technology to understand the various implementations disclosed herein.
Claims
1. A method for controlling virtual objects in an application, comprising: in response to determining that a first action sequence of a user of the application matches a first predetermined action sequence, presenting a set of virtual objects in a display area of the application in a first format; determining a target virtual object among the set of virtual objects based on user interaction of the user; and in response to determining that a second action sequence of the user matches a second predetermined action sequence, performing an operation corresponding to the target virtual object.
2. The method according to claim 1, further comprising: presenting the first action sequence in the display area, wherein a presentation position of the set of virtual objects in the display area corresponds to a presentation position of the first action sequence in the display area.
3. The method according to claim 1, further comprising: presenting the target virtual object in a second format different from the first format in the display area.
4. The method according to claim 3, further comprising: presenting the selected target virtual object in a third format different from the first format and the second format in the display area.
5. The method according to claim 1, further comprising: in response to detecting a third action sequence of the user, performing the operation corresponding to the target virtual object based on the third action sequence.
6. The method according to claim 1, further comprising: in response to determining that a fourth action sequence of the user matches a third predetermined action sequence, exiting the operation corresponding to the target virtual object.
7. The method according to claim 6, wherein the third predetermined action sequence includes the first predetermined action sequence.
8. The method according to claim 3, further comprising at least any one of the following: in response to determining that the user interaction matches another target virtual object among the set of virtual objects, presenting the another target virtual object in the second format in the display area; and in response to not detecting an action sequence matching the second predetermined action sequence within a predetermined time period and the user interaction not matching any virtual object among the set of virtual objects, removing the set of virtual objects from the display area.
9. The method according to claim 1, wherein the predetermined action sequences in the first predetermined action sequence and the second predetermined action sequence are specified by the user.
10. The method according to claim 9, further comprising: in response to detecting user interaction for specifying the predetermined action sequence, receiving a video for specifying the predetermined action sequence; and determining the predetermined action sequence from the video.
11. The method according to claim 1, wherein any one of the first predetermined action sequence and the second predetermined action sequence includes an action sequence performed by any one of the following: an arm, both hands, and a single hand, and an end action in the first predetermined action sequence is the same as a start action in the second predetermined action sequence.
12. The method according to claim 1, wherein: Any one of the first action sequence and the second action sequence is determined based on data collected from at least any one of the following: an environmental sensor in a wearable device worn by the user, and an accessory device associated with the wearable device; The line-of-sight direction of the user interaction is determined based on data collected by an eye-tracking sensor in the wearable device; and The application runs on a computing device in the wearable device.
13. The method according to claim 12, wherein the action sequence includes at least one action, and the action includes at least any one of the following: An interaction action of the user with the accessory device; and A body action of at least one body part of the user, and the body action is represented by at least any one of the following: the posture and position of at least one body part of the user.
14. The method according to claim 1, wherein the user interaction includes user interaction based on at least any one of the following: the line-of-sight direction of the user, the eye movement of the user, the voice interaction of the user, and the limb interaction of the user.
15. A device for controlling a virtual object in an application, comprising: A first presentation module configured to present a set of virtual objects in a first format in a display area of the application in response to determining that a first action sequence of a user of the application matches a first predetermined action sequence; A determination module configured to determine a target virtual object in the set of virtual objects based on the user interaction of the user; and An execution module configured to perform an operation corresponding to the target virtual object in response to determining that a second action sequence of the user matches a second predetermined action sequence.
16. An electronic device, comprising: At least one processing unit; and At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processing unit.
17. A computer-readable storage medium having stored thereon a computer program, which when executed by a processor causes the processor to implement the method according to any one of claims 1 to 14.