Somatosensory intelligent terminal interaction method, device, equipment and storage medium
The user's action data is obtained through the camera and grid processing is performed, and the relative movement trend is determined to generate operation instructions, which solves the problem that users find it difficult to control smart terminals when inconvenience is achieved, touch-free interaction, and improves the user experience.
Patent Information
- Application Number
- CN202210419153.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-04-20
AI Technical Summary
The existing smart terminal interaction method is difficult to accurately control the smart terminal when there is water or stains on the user's hands, or when elderly users are not flexible enough in their fingers and have poor vision, resulting in poor user experience.
The user's action data is obtained through the camera of the somatosensory intelligent terminal, and grid processing is performed, the similarity to the preset image is compared, and the relative movement trend of the characteristic objects is determined, thereby generating operation instructions and controlling the operation of the intelligent terminal.
It realizes the function of users controlling smart terminals without touching the screen, and improves the user's interactive experience, especially in hand inconvenience or use scenarios for elderly users.
Smart Images

Figure CN114816057B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of human-computer interaction technology, and in particular, to a method, device, equipment and storage medium for somatosensory intelligent terminal interaction. Background Art
[0002] With the rapid development of the Internet of Things and 5G technology, people's living standards are constantly improving. In order to pursue high-efficiency and high-quality life, intelligent technology and somatosensory technology are widely used in daily life. The existing interaction method requires touch operation to control the device, but in some application scenarios, it is not convenient for people to use their fingers to touch the smart terminal. For example, when the hands are stained with water or stains, it is difficult for people to use touch to accurately control the smart terminal. In addition, when the user is an elderly person, due to physical reasons such as insufficient finger flexibility and poor eyesight, it is even more difficult to accurately operate the device, resulting in a poor user experience. Summary of the invention
[0003] The embodiments of the present application provide a somatosensory intelligent terminal interaction method, device, equipment and storage medium, which can provide users with effective solutions when it is inconvenient for users to perform touch operations, thereby improving the user experience.
[0004] In a first aspect, an embodiment of the present application provides a method for interacting with a somatosensory intelligent terminal, the method comprising:
[0005] Acquire first motion data of a user through a camera of a somatosensory intelligent terminal, where the first motion data includes multiple frames of image data;
[0006] Obtaining a grid-covered image of the first frame of image data by performing grid processing on the first frame of image data in the first action data;
[0007] Confirm whether there is a first feature image in the preset images in the action information library whose similarity with the grid-covered image is greater than a first threshold;
[0008] If the first characteristic image exists, performing the gridding process on the remaining image data to obtain a grid-covered image of the remaining image data;
[0009] Determining a relative movement trend of a feature object in the first motion data according to a comparison result between the plurality of grid-covered images;
[0010] According to the relative movement trend, an operation instruction for the somatosensory intelligent terminal is determined.
[0011] In a second aspect, an embodiment of the present application further provides a somatosensory intelligent terminal interaction device, the device comprising:
[0012] A first acquisition module is configured to acquire first motion data of a user through a camera of a somatosensory intelligent terminal, wherein the first motion data includes multiple frames of image data;
[0013] A second acquisition module is configured to obtain a grid-covered image of the first frame of image data by performing grid processing on the first frame of image data in the first action data;
[0014] A similarity judgment module is configured to confirm whether there is a first feature image in the preset images in the action information library whose similarity with the grid cover image is greater than a first threshold;
[0015] a response module, configured to, if the first feature image exists, perform the gridding process on the remaining image data to obtain a grid-covered image of the remaining image data;
[0016] A first determination module, configured to determine a relative movement trend of a feature object in the first motion data according to a comparison result between the plurality of grid-covered images;
[0017] The second determination module is configured to determine an operation instruction for the somatosensory intelligent terminal according to the relative movement trend.
[0018] In a third aspect, an embodiment of the present application further provides a computer device, the device comprising:
[0019] one or more processors;
[0020] A storage device for storing one or more programs;
[0021] When one or more of the programs are executed by one or more of the processors, the one or more processors implement the somatosensory intelligent terminal interaction method described in the embodiment of the present application.
[0022] In a fourth aspect, an embodiment of the present application further provides a storage medium storing computer executable instructions, which, when executed by a computer processor, are used to execute the somatosensory intelligent terminal interaction method described in the embodiment of the present application.
[0023] In an embodiment of the present application, by obtaining the user's first action data, such as video data, etc., multiple frames of images therein are gridded in succession to obtain a grid coverage image, and it is possible to determine whether to respond based on the similarity between the grid coverage image and the preset image, and further determine the relative movement trend based on the comparison between the multiple grid coverage images, and then determine the operation instructions for the somatosensory intelligent terminal, so as to achieve the effect of controlling the somatosensory intelligent terminal. That is, from the user's point of view, the display content of the terminal display screen can be controlled through their own actions, without touching the display screen to control it, thereby improving the user's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A flowchart of a somatosensory intelligent terminal interaction method provided in an embodiment of the present application;
[0025] Figure 2 A schematic diagram of a somatosensory intelligent terminal communication system provided in an embodiment of the present application;
[0026] Figure 3 A structural block diagram of a somatosensory intelligent terminal interaction device provided in an embodiment of the present application;
[0027] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, rather than to limit the embodiments of the present application. It should also be noted that, for ease of description, only parts related to the embodiments of the present application are shown in the accompanying drawings, rather than all structures.
[0029] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship. Among them, "several" means one or more, and "multiple" means two or more.
[0030] Figure 1 The flowchart of the somatosensory intelligent terminal interaction method provided in the embodiment of the present application can be used in the process of interaction. The method can be executed by a computing device for somatosensory analysis such as a somatosensory intelligent terminal, a server, etc., and specifically includes the following steps:
[0031] Step S100: Acquire the user's first motion data through the camera of the somatosensory intelligent terminal.
[0032] The somatosensing intelligent terminal can be a terminal device such as a mobile phone with a camera, a tablet computer with a camera, etc., which obtains the user's first action data through the camera thereon. The first action data includes multiple frames of image data, that is, the user's actions in front of the camera are captured by the camera, such as recording in the form of video, thereby decomposing multiple frames of image data as the first action data.
[0033] Step S200 , obtaining a grid-covered image of the first frame of image data by gridding the first frame of image data in the first action data.
[0034] A first frame of image data in the multiple frames of image data of the first action data is gridded, and the first frame of image data is used as a starting frame in the multiple frames of image data, and is gridded to obtain a grid covered image corresponding to the first frame of image data.
[0035] In one embodiment, the image data is gridded, and a Delaunay triangulation algorithm may be used to generate a Voronoi diagram corresponding to the image data. The image data is grid-covered according to the Voronoi diagram, thereby generating a grid-covered image.
[0036] It is understandable that for image data, it is necessary to perform Delaunay triangulation on the image so as to determine the base points on the image, and the dividing lines on the Voronoi diagram can be determined through the base points, wherein the dividing lines are the perpendicular bisectors of two adjacent base points. It is conceivable that in the Voronoi diagram, the plane is divided into multiple regions by the dividing lines, and the base points are located in the above regions. Therefore, when the Voronoi diagram is overlaid on the image, a grid-covered image can be obtained, that is, the grid-covered image has the image data and the base points and dividing lines of the Voronoi diagram.
[0037] It should be conceivable that partial positions of an image in the image data may be gridded. For example, if the image data includes an image of a gesture made by a user in front of a camera, the user's hand may be taken as a feature object and gridded.
[0038] Step S300: confirm whether there is a first feature image in the preset images in the action information library whose similarity with the grid cover image is greater than a first threshold.
[0039] The action information library is stored in a memory and includes multiple preset images, which are also gridded images. It is determined whether there is a first feature image in the multiple preset images whose similarity with the grid-covered image of the first frame image data is above a first threshold. The first feature image is one of the multiple preset images, and its similarity with the grid-covered image is greater than the first threshold.
[0040] For example, in an application scenario, corresponding to a certain instruction of the user, the preset image is an image of a certain gesture made by the user. When the gesture on the grid overlay image is similar to the gesture in the preset image, and the similarity is above a first threshold, it can be determined that the preset image corresponds to the network overlay image, and then subsequent operations are performed. The first threshold can be set according to the design accuracy, such as 80%.
[0041] In one embodiment, the confirmation of the first feature image can be determined by judging the similarity between the grid covering image and the preset image. It can be understood that there are several preset images in the action information library, and it can be determined by traversing whether the similarity between the preset image and the grid covering image is greater than the first threshold, thereby determining that the preset image is the first feature image.
[0042] For example, base points determined after gridding of a grid-covered image and base points determined after gridding of a preset image are obtained, and the number of overlapping base points between the two is compared, where the number of overlapping base points can be in the form of a percentage, for determining the proportion of the overlapping base points in the total number of base points of the grid-covered image, thereby determining the similarity between the grid-covered image and the preset image; when the similarity between the grid-covered image and the preset image is greater than a first threshold, the preset image is used as a first feature image, that is, it can be determined that the first feature image exists in several preset images.
[0043] Step S400: If the first characteristic image exists, gridding is performed on the remaining image data to obtain a grid coverage image of the remaining image data.
[0044] It can be understood that, in the presence of the first feature image, the remaining image data in the first action data is gridded, such as using the Delaunay triangulation algorithm to generate a Voronoi diagram corresponding to the image data, and the image data is grid-covered according to the Voronoi diagram, thereby obtaining a grid-covered image corresponding to the remaining image data.
[0045] Step S500: Determine the relative movement trend of the feature object in the first motion data according to the comparison results between the multiple grid coverage images.
[0046] It can be understood that each frame of image data in the first action data corresponds to a grid coverage image. By comparing the grid coverage images of two adjacent frames of image data, it should be considered that the two adjacent frames of image data are forward compared, such as the grid coverage image of the second frame of image data and the grid coverage image of the first frame of image data, the grid coverage image of the third frame of image data and the grid coverage image of the second frame of image data, until the grid coverage image of the last frame of image data.
[0047] After completing the comparison between multiple frames of image data, the relative movement trend of the feature objects in the image can be determined. It can be imagined that in an application scenario, the camera captures multiple frames of images of the user performing gestures, and the user's hands in the image can be used as feature objects, such as using an image recognition algorithm to identify the hands in the image. It should be noted that in some embodiments, only the hand area can be gridded.
[0048] Step S600: Determine an operation instruction for the somatosensory intelligent terminal according to the relative movement trend.
[0049] After the relative movement trend of the characteristic object is determined, the operation instruction for the somatosensory intelligent terminal can be determined according to the relative movement trend.
[0050] In one embodiment, when the relative movement trend is a first direction, the operation instruction for the somatosensing intelligent terminal is determined to be a first scrolling display instruction; when the relative movement trend is a second direction, the operation instruction for the somatosensing intelligent terminal is determined to be a second scrolling display instruction; when the relative movement trend is a third direction, the operation instruction for the somatosensing intelligent terminal is determined to be an entry instruction; when the relative movement trend is a fourth direction, the operation instruction for the somatosensing intelligent terminal is determined to be an exit instruction.
[0051] Exemplarily, in an application scenario, if the somatosensing intelligent terminal is a mobile phone, when the user browses web content by using the mobile phone, the user can control the somatosensing intelligent terminal through gestures, such as the user makes a gesture in front of the front camera of the mobile phone. If the gesture is a preset gesture, such as the preset gesture is five fingers together, corresponding to the user maintaining the gesture and moving upward, that is, the relative movement trend is a first direction, the operation instruction of the somatosensing intelligent terminal is determined to be a first scrolling display instruction, that is, the somatosensing intelligent terminal controls the display content of the display screen to scroll upward according to the first scrolling display instruction.
[0052] Corresponding to the user maintaining the gesture and moving downward, that is, the relative movement trend is in the second direction, the determined operation instruction of the somatosensory intelligent terminal is the second scrolling display instruction, that is, the somatosensory intelligent terminal controls the display content of the display screen to scroll downward according to the second scrolling display instruction.
[0053] Corresponding to the user maintaining the gesture and moving to the left, that is, the relative movement trend is the third direction, the operation instruction of the somatosensory intelligent terminal is determined to be an entry instruction, that is, the somatosensory intelligent terminal enters the next level link of the web page according to the entry instruction, such as the next page, and the display screen updates the display content.
[0054] Corresponding to the user maintaining the gesture and moving to the right, that is, the relative movement trend is the fourth direction, the operation instruction of the somatosensing intelligent terminal is determined to be an exit instruction, that is, the somatosensing intelligent terminal exits the current web page or returns to the previous level link, such as the previous page, according to the exit instruction, and the display screen updates the display content.
[0055] It should be understood that movement up, down, left and right refers to the movement of the user's hand on a certain plane within the shooting range of the camera, such as the user's hand moving on a plane parallel to the camera, and the four directions perpendicular to each other on the plane are used as the up, down, left and right directions.
[0056] It can be seen from the above scheme that the embodiment of the present application obtains the user's first action data, such as video data, etc., and grids the first frame of image data therein to obtain its grid coverage image, so as to determine whether to respond based on the similarity between the grid coverage image and the preset image. When the first feature image exists, the remaining image data is gridded, so that the relative movement trend can be determined based on the comparison between multiple grid coverage images, and then the operation instructions for the somatosensory intelligent terminal can be determined, so as to achieve the effect of controlling the somatosensory intelligent terminal. That is, from the user's point of view, the somatosensory intelligent terminal can be controlled through his own actions without touch operation, so that the user's interactive experience is improved.
[0057] In one embodiment, when the first feature image does not exist in the action information library, that is, there is no preset image among the multiple preset images in the action information library whose similarity with the grid-covered image of the first frame image data is greater than the first threshold, but if there is a second feature image therein, the second feature image is a preset image whose similarity with the grid-covered image of the first frame image data is greater than the second threshold, wherein the second threshold can be set according to the design accuracy, such as being set to 60%.
[0058] In the presence of a second feature image, the grid overlay image of the first frame image data is added to the action information library. It can be understood that in the action information library, one or more preset images correspond to an operation instruction. Therefore, the number of preset images corresponding to the operation instruction is increased, thereby making it easier to trigger the somatosensory intelligent terminal to perform corresponding operations according to the operation instruction, thereby improving the user experience.
[0059] In one embodiment, the preset images in the action information library can also be expanded. The second action data is obtained through a camera, wherein the second action data includes a plurality of frames of image data. The plurality of frames of image data in the second action data are all gridded to obtain a plurality of grid-covered images corresponding to the plurality of frames of image data in the second action data, i.e., preset images. The plurality of preset images are stored in the action information library to expand the preset images in the action information library.
[0060] Exemplarily, in the action information library, one or more preset images correspond to an operation instruction; the user can define the content of the action information library according to the preset settings, such as defining the image data corresponding to the operation instruction or the operation instruction corresponding to the image data, etc. When the user's hand is used as the feature object, the user can use the camera to capture the image of his gesture action, use the image of the gesture action as image data to define a certain operation instruction or a new operation instruction, and grid the image data for storage in the action information library.
[0061] From the above scheme, it can be seen that users can use the somatosensory intelligent terminal to enter the preset images in the action information library, that is, users can define the actions of feature objects in the preset images stored in the action information library according to their own operating habits, so that users can more flexibly configure interactive actions and corresponding operating instructions, thereby improving the user experience.
[0062] In some embodiments, for the grid cover image corresponding to each frame of image data and the grid cover image corresponding to the previous frame of image data, the inter-frame movement trend is obtained, and the inter-frame movement trend is used to represent the movement trend of the feature object between two adjacent frames. For the inter-frame movement trend, the memory has a preset trend proportion, so the relative movement trend can be determined according to the preset trend proportion. It should be conceived that the preset trend proportion is set with a weight value corresponding to each inter-frame movement trend, and each weight value in the preset trend proportion is successively reduced, and the weight value of the inter-frame movement trend between the first frame of image data and the second frame of image data is the largest.
[0063] Exemplarily, there are 5 frames of image data in the first action data. For example, the grid lines in the grid overlay image I corresponding to the first frame of image data are located at position A in the image, and the grid lines in the grid overlay image II in the second frame of image data are located at position B in the image. The two grid overlay images overlap, and it can be determined that position B is above position A, that is, the grid line can reach position B by moving upward at position A, that is, the inter-frame movement trend is upward. It can be imagined that the grid line from position A to position B can be quantified into a vector, and the vector can be decomposed into two movement components corresponding to the up and down directions and the left and right directions. Based on the size of the movement component, the largest movement component is selected and its direction is used as the inter-frame movement trend.
[0064] Therefore, for the inter-frame movement trend between the 5 frames of image data, inter-frame movement trend a, inter-frame movement trend b, inter-frame movement trend c and inter-frame movement trend d can be obtained. The preset trend proportion can be set to a:b:c:d=40%:30%:20%:10%. If the inter-frame movement trend a is upward, the inter-frame movement trend b is upward, the inter-frame movement trend c is left and the inter-frame movement trend d is left, it can be determined that the upward inter-frame movement trend accounts for a larger proportion, thereby determining that the relative movement trend is upward.
[0065] It can be seen from the above scheme that for the relative movement trend of the feature object in the first action data, the inter-frame movement trend between each grid-covered image is taken as a basis, and a preset trend proportion is also set for each inter-frame movement trend. The proportion of the inter-frame movement trend to the relative movement trend is determined by different weight ratios. By setting successively decreasing weight values, the inter-frame movement trend corresponding to the first frame image data and the second frame image data occupies a larger proportion, which can effectively determine the movement direction of the feature object and reduce the interference of subsequent image data on the relative movement trend due to human factors.
[0066] The somatosensory intelligent terminal interaction method in the embodiment of the present application can be implemented based on an intelligent terminal such as a mobile phone, that is, the data processing process in step S100-step S600 is completed by the intelligent terminal. In the specific implementation, in order to reduce the data processing burden of the intelligent terminal, the somatosensory intelligent terminal interaction method in the embodiment of the present application can be further implemented through the collaborative processing and cooperation of multiple physical devices according to the network environment where the intelligent terminal is located.
[0067] Figure 2 A schematic diagram of a somatosensory intelligent terminal communication system provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the system includes a somatosensory intelligent terminal 201, a 5G CPE (Customer Premise Equipment) terminal 202, a 5G base station 203, and a somatosensory analysis server 204. The somatosensory intelligent terminal 201 can be a mobile phone, a flat screen with a camera, etc. The somatosensory intelligent terminal 201 establishes a wired or wireless connection with the 5G CPE terminal 202. For example, when the somatosensory intelligent terminal 201 is a mobile phone, it is connected to the 5G CPE terminal 202 wirelessly. The 5G CPE terminal 202 is connected to the 5G base station 203 for communication. In addition, the somatosensory analysis server 204 is connected to the 5G base station 203 for communication. The somatosensory analysis server 204 is used to execute the somatosensory intelligent terminal interaction method described in the above embodiment.
[0068] In one application scenario, a user uses a somatosensory intelligent terminal 201 such as a mobile phone to browse the web, and uses the hand as a feature object. The user makes a gesture in front of the camera, and the mobile phone obtains multiple frames of image data of the user's gesture through the camera, that is, the first motion data, and transmits the first motion data to the 5G base station 203 device through the 5G CPE terminal 202. After receiving the first motion data, the somatosensory analysis server 204 grids the first frame of image data in the first motion data to obtain a grid coverage image of the first frame of image data. When the similarity of the grid coverage image of the first frame of image data is greater than a first threshold, the somatosensory analysis server 204 further grids the remaining image data to obtain a grid coverage image of the remaining image data, thereby comparing the grid coverage image to determine the relative movement trend of the user's hand in the first motion data, and then determine the operation instructions for the mobile phone.
[0069] The body sensing analysis server 204 sends the operation instruction through the 5G base station 203, and the 5G base station 203 transmits the operation instruction to the user's mobile phone through the 5G CPE terminal 202. After the mobile phone receives the operation instruction, the mobile phone will respond according to the operation instruction. If the operation instruction is an exit instruction, the mobile phone will exit the current web page.
[0070] It can be seen from the above scheme that the somatosensory intelligent terminal communication system can provide users with a contactless operation experience, allowing users to control the somatosensory intelligent terminal more conveniently. In some scenarios where it is inconvenient for both hands to touch the terminal device, such as when the user's hands are wet, at this time, the user uses his fingers to touch the mobile phone, which will cause the mobile phone screen to be stained with water droplets, resulting in insensitive touch and difficulty in controlling the mobile phone. Through the scheme of this application, the user can control the mobile phone through gestures, such as controlling the mobile phone to exit the current page and other operations. Therefore, this application can solve the problem of users not being able to use both hands to touch the operation, and effectively provide users with a better solution, so that users can effectively and conveniently control the mobile phone, thereby improving the user experience.
[0071] Moreover, in some application scenarios, such as when the elderly use smart terminals, it is difficult for the elderly to accurately use their fingers to click on the correct touch controls due to their inflexible fingers and poor eyesight. In the application of the present application, some gestures can be pre-set to achieve control of the smart terminal, so that when the elderly use the smart terminal, they can complete partial control of the smart terminal by making gestures and moving their hands, effectively reducing the difficulty for the elderly to use smart terminals, thereby better taking into account the needs of the user groups and providing users with a better interactive experience.
[0072] Figure 3This is a structural block diagram of a somatosensory intelligent terminal interaction device provided in an embodiment of the present application. The device is used to execute the somatosensory intelligent terminal interaction method provided in the above embodiment, and has functional modules and beneficial effects corresponding to the execution method, such as Figure 3 As shown, the device specifically includes: a first acquisition module 301, a first acquisition module 302, a similarity judgment module 303, a response module 304, a first determination module 305 and a second determination module 306, wherein:
[0073] A first acquisition module 301 is configured to acquire first motion data of a user through a camera of a somatosensory intelligent terminal, wherein the first motion data includes multiple frames of image data;
[0074] A second acquisition module 302 is configured to obtain a grid-covered image of the first frame of image data by performing grid processing on the first frame of image data in the first action data;
[0075] A similarity determination module 303 is configured to determine whether there is a first feature image in the preset images in the action information library whose similarity with the grid cover image is greater than a first threshold;
[0076] The response module 304 is configured to perform the gridding process on the remaining image data to obtain a grid-covered image of the remaining image data if the first feature image exists;
[0077] A first determination module 305 is configured to determine a relative movement trend of a feature object in the first motion data according to a comparison result between the plurality of grid-covered images;
[0078] The second determination module 306 is configured to determine an operation instruction for the somatosensory intelligent terminal according to the relative movement trend.
[0079] It can be seen from the above scheme that in the embodiment of the present application, the user's first action data is acquired by the first acquisition module, and the second acquisition module performs grid processing on the first frame image data therein to obtain a grid coverage image. The similarity judgment module can determine whether there is a first feature image based on the similarity between the grid coverage image and the preset image. If the first feature image exists, the response module performs grid processing on the remaining image data, and the first determination module also determines the relative movement trend based on the comparison between multiple grid coverage images, so that the second determination module determines the operation instructions of the somatosensory intelligent terminal based on the relative movement trend, thereby achieving the effect of controlling the somatosensory intelligent terminal, that is, from the user's point of view, the somatosensory intelligent terminal can be controlled through his own actions, such as controlling the display screen to update the display content without touching its display screen, thereby improving the user's interactive experience.
[0080] In one embodiment, the similarity judgment module 303 is configured to confirm whether there is a second feature image in the preset image whose similarity with the grid coverage image is greater than a second threshold if the first feature image does not exist in the preset image; if the second feature image exists, add the grid coverage image of the first frame image data to the action information library.
[0081] In some embodiments, the first acquisition module 301 is configured to acquire second motion data preset by the user through the camera of the somatosensory intelligent terminal, where the second motion data includes several frames of image data; by gridding several frames of image data in the second motion data, several preset images are obtained, and the several preset images are stored in the motion information library.
[0082] In some embodiments, the similarity judgment module 303 is configured to obtain the base points determined by the grid-covered image and the preset image based on the gridding processing; determine the similarity between the grid-covered image and the preset image according to the number of overlapping base points between the grid-covered image and the preset image, and confirm the preset image whose similarity is greater than a first threshold as the first feature image.
[0083] In some embodiments, the second acquisition module 302 is configured to generate a Voronoi diagram corresponding to the image data based on a Delaunay triangulation algorithm; and perform grid coverage on the image data according to the Voronoi diagram to generate a grid coverage image.
[0084] In some embodiments, the first determination module 305 is configured to compare the grid coverage image corresponding to each frame of image data with the grid coverage image corresponding to the previous frame of image data to obtain the inter-frame movement trend; and determine the relative movement trend based on a preset trend proportion of the inter-frame movement trend.
[0085] In some embodiments, the second determination module 306 is configured to, when the relative movement trend is a first direction, determine that the operation instruction to the somatosensing intelligent terminal is a first scrolling display instruction; when the relative movement trend is a second direction, determine that the operation instruction to the somatosensing intelligent terminal is a second scrolling display instruction; when the relative movement trend is a third direction, determine that the operation instruction to the somatosensing intelligent terminal is an entry instruction; when the relative movement trend is a fourth direction, determine that the operation instruction to the somatosensing intelligent terminal is an exit instruction.
[0086] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application is shown in FIG. Figure 4As shown, the device includes a processor 401, a memory 402, an input device 403 and an output device 404; the number of processors 401 in the device can be one or more. Figure 4 A processor 401 is taken as an example; the processor 401, memory 402, input device 403 and output device 404 in the device can be connected by a bus or other means. Figure 4 The example of connecting through a bus is taken. The memory 402, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the somatosensory intelligent terminal interaction method in the embodiment of the present application. The processor 401 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 402, that is, realizes the above-mentioned somatosensory intelligent terminal interaction method. The input device 403 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 404 may include a display device such as a display screen.
[0087] The embodiment of the present application further provides a storage medium containing computer executable instructions, wherein the computer executable instructions are used to execute a somatosensory intelligent terminal interaction method described in the above embodiment when executed by a computer processor, specifically including:
[0088] Acquire first motion data of a user through a camera of a somatosensory intelligent terminal, where the first motion data includes multiple frames of image data;
[0089] Obtaining a grid-covered image of the first frame of image data by performing grid processing on the first frame of image data in the first action data;
[0090] Confirm whether there is a first feature image in the preset images in the action information library whose similarity with the grid-covered image is greater than a first threshold;
[0091] If the first characteristic image exists, performing the gridding process on the remaining image data to obtain a grid-covered image of the remaining image data;
[0092] Determining a relative movement trend of a feature object in the first motion data according to a comparison result between the plurality of grid-covered images;
[0093] According to the relative movement trend, an operation instruction for the somatosensory intelligent terminal is determined.
[0094] Storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0095] Note that the above are only preferred embodiments of the present application and the technical principles used. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A somatosensory intelligent terminal interaction method, characterized in that: include: Acquire first motion data of a user through a camera of a somatosensory intelligent terminal, where the first motion data includes multiple frames of image data; Obtaining a grid-covered image of the first frame of image data by performing grid processing on the first frame of image data in the first action data; Confirm whether there is a first feature image in the preset images in the action information library whose similarity with the grid-covered image is greater than a first threshold; If the first feature image exists, performing the gridding process on the remaining image data to obtain a grid-covered image of the remaining image data; Determining a relative movement trend of a feature object in the first motion data according to a comparison result between the plurality of grid-covered images; Determining an operation instruction for the somatosensory intelligent terminal according to the relative movement trend; The gridding process includes: Based on the Delaunay triangulation algorithm, generating a Voronoi diagram corresponding to the image data; Performing grid coverage on the image data according to the Voronoi diagram to generate a grid coverage image; Determining the relative movement trend of the feature object in the first motion data according to the comparison results between the plurality of grid covered images includes: Compare the grid coverage image corresponding to each frame of image data with the grid coverage image corresponding to the previous frame of image data to obtain the inter-frame movement trend; The relative movement trend is determined according to a preset trend proportion of the inter-frame movement trend.
2. The method according to claim 1, characterized in that After confirming whether there is a first feature image in the preset images in the action information library whose similarity with the grid cover image is greater than a first threshold, the method further includes: If the first characteristic image does not exist in the preset image, confirming whether a second characteristic image having a similarity with the grid coverage image greater than a second threshold value exists in the preset image; If the second feature image exists, a grid overlay image of the first frame image data is added to the action information library.
3. The method according to claim 1, characterized in that Before obtaining the first action data of the user, the method further includes: Acquiring second motion data preset by a user through a camera of the somatosensory intelligent terminal, wherein the second motion data includes a plurality of frames of image data; By performing gridding processing on a plurality of frames of image data in the second action data, a plurality of the preset images are obtained, and the plurality of the preset images are stored in the action information library.
4. The method according to claim 1, characterized in that: The step of confirming whether there is a first feature image in the preset images in the action information library whose similarity with the grid coverage image is greater than a first threshold comprises: Acquire the grid covered image and the base points of the preset image determined based on the gridding process; The similarity between the grid-covered image and the preset image is determined according to the number of overlapping base points between the grid-covered image and the preset image, and the preset image whose similarity is greater than a first threshold is confirmed as the first feature image.
5. The method according to claim 1, characterized in that The step of determining the operation instruction for the somatosensory intelligent terminal according to the relative movement trend includes: When the relative movement trend is a first direction, determining that the operation instruction for the somatosensory intelligent terminal is a first scrolling display instruction; When the relative movement trend is the second direction, determining that the operation instruction for the somatosensory intelligent terminal is a second scrolling display instruction; When the relative movement trend is a third direction, determining that the operation instruction to the somatosensory intelligent terminal is an entry instruction; When the relative movement trend is the fourth direction, it is determined that the operation instruction to the somatosensory intelligent terminal is an exit instruction.
6. A somatosensory intelligent terminal interaction device, characterized in that: include: A first acquisition module is configured to acquire first motion data of a user through a camera of a somatosensory intelligent terminal, wherein the first motion data includes multiple frames of image data; A second acquisition module is configured to obtain a grid-covered image of the first frame of image data by performing grid processing on the first frame of image data in the first action data; A similarity judgment module is configured to confirm whether there is a first feature image in the preset images in the action information library whose similarity with the grid cover image is greater than a first threshold; a response module, configured to, if the first feature image exists, perform the gridding process on the remaining image data to obtain a grid-covered image of the remaining image data; A first determination module, configured to determine a relative movement trend of a feature object in the first motion data according to a comparison result between the plurality of grid-covered images; A second determination module is configured to determine an operation instruction for the somatosensory intelligent terminal according to the relative movement trend; The gridding process includes: Based on the Delaunay triangulation algorithm, generating a Voronoi diagram corresponding to the image data; Performing grid coverage on the image data according to the Voronoi diagram to generate a grid coverage image; The first determination module is specifically configured as follows: Compare the grid coverage image corresponding to each frame of image data with the grid coverage image corresponding to the previous frame of image data to obtain the inter-frame movement trend; The relative movement trend is determined according to a preset trend proportion of the inter-frame movement trend.
7. A computer device, characterized in that: The device comprises: one or more processors; A storage device for storing one or more programs; When one or more of the programs are executed by one or more of the processors, the one or more processors implement the somatosensory intelligent terminal interaction method as described in any one of claims 1-5.
8. A storage medium storing computer executable instructions, characterized in that: The computer executable instructions are used to execute the somatosensory intelligent terminal interaction method as described in any one of claims 1 to 5 when executed by a computer processor.
Citation Information
Patent Citations
Gesture recognition method and device
CN112364799A
A gesture-based man-machine interaction method and equipment
CN112965602A
Gesture recognition method, device and system and vehicle
CN113646736A