A gesture interaction control method, device, equipment, storage medium
By detecting hand gestures in the cloud film system, acquiring and processing hand motion video streams, determining hand motion tags, and mapping controls, the problem of inconvenient operation of the cloud film system is solved, realizing convenient and vivid interaction without mouse operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2026-04-14
AI Technical Summary
When conducting consultations or teaching on a large screen, the existing cloud film system requires the use of a mouse, which is inconvenient and difficult to understand.
By detecting when the cloud film system has started gesture recognition, the system acquires video streams of hand movements, performs video frame extraction and image processing, determines the hand movement labels of the target video frames, and then performs interactive control of the cloud film system based on the mapping relationship.
This allows users to operate the cloud film system freely without the need for a mouse, improving the convenience and vividness of operation and enhancing communication efficiency.
Smart Images

Figure CN115993887B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a gesture interaction control method, device, equipment, and storage medium. Background Technology
[0002] With the continuous development of software technology, many web-based medical image browsing front-ends have emerged in the cloud imaging industry, such as the open-source ohif. The existing technologies are all based on cornerstone technology, which can display images well, but there are the following problems: when consultations or teaching are needed on a large screen, everyone needs to go to a designated location to operate the cloud film system with a mouse, which is inconvenient and difficult to understand.
[0003] In summary, how to achieve free operation of the cloud film system without the aid of external devices such as a mouse is a technical problem that needs to be solved in this field. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a gesture-based interactive control method, device, equipment, and storage medium that enables free operation of a cloud film system without the aid of external devices such as a mouse. The specific solution is as follows:
[0005] In a first aspect, this application discloses a gesture interaction control method, characterized in that it is applied to a cloud film system and includes:
[0006] When the cloud film system detects that gesture recognition is enabled, a video stream of hand movements is captured.
[0007] Perform video frame extraction on the video stream to determine the target video frame;
[0008] The corresponding hand gesture tag for the target video frame is determined, and the cloud film system is interactively controlled based on the hand gesture tag and the cloud film system control mapping relationship.
[0009] Optionally, the step of performing video frame extraction on the video stream to determine the target video frame includes:
[0010] The video stream is subjected to video frame extraction and image conversion processing to determine the target video frame.
[0011] Optionally, the step of performing video frame extraction and image conversion processing on the video stream includes:
[0012] Perform video frame extraction on the video stream to obtain the video frames to be processed;
[0013] The video frames to be processed are sequentially subjected to bilateral filtering and mirror flipping to obtain preprocessed video frames.
[0014] Optionally, determining the target video frame includes:
[0015] The preprocessed video frame is cropped to obtain a rectangular region containing the outline of the hand, and the rectangular region is used as the target video frame.
[0016] Optionally, determining the target video frame includes:
[0017] The rectangular region is subjected to background removal processing, and the processed rectangular region is then subjected to image grayscale processing, filtering processing, and binarization processing in sequence to determine the target video frame.
[0018] Optionally, determining the corresponding hand gesture tag for the target video frame includes:
[0019] Obtain the hand indentation of the target video frame, and determine the number of finger pits based on the hand indentation;
[0020] The target video frame is tagged based on the number of finger sockets to determine the corresponding hand action tag for the target video frame.
[0021] Optionally, the interactive control of the cloud film system based on the hand gesture tag and the cloud film system control mapping relationship includes:
[0022] The hand movement trajectory is tracked based on the hand movement tag of each target video frame, and the correspondence between the coordinate points of the finger in the target video frame and the target coordinate points corresponding to the control operation of the cloud film system is determined based on the hand movement trajectory.
[0023] Based on the aforementioned correspondence, the cloud film system is controlled to complete the interactive control of the cloud film system.
[0024] Secondly, this application discloses a gesture interaction control device, characterized in that it is applied to a cloud film system and includes:
[0025] The video stream acquisition module is used to acquire a video stream of hand movements when the cloud film system detects that gesture recognition operation has been enabled.
[0026] The video frame extraction module is used to perform video frame extraction operations on the video stream to determine the target video frame;
[0027] An interactive control module is used to determine the corresponding hand gesture tag of the target video frame and to perform interactive control of the cloud film system based on the hand gesture tag and the cloud film system control mapping relationship.
[0028] Thirdly, this application discloses an electronic device, including:
[0029] Memory, used to store computer programs;
[0030] A processor is configured to execute the computer program to implement the steps of the aforementioned disclosed gesture interaction control method.
[0031] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed gesture interaction control method.
[0032] Therefore, this application discloses a method for acquiring a video stream of hand movements when the cloud film system is activated for gesture recognition; performing video frame segmentation on the video stream to determine a target video frame; determining the corresponding hand movement tag for the target video frame; and interactively controlling the cloud film system based on the mapping relationship between the hand movement tag and the cloud film system control. It is evident that operating the cloud film system through gesture recognition facilitates efficient communication for doctors. Gesture operation allows for more vivid and intuitive manipulation of the cloud film system, replacing traditional mouse operation with user-friendly gestures and offering greater freedom. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0034] Figure 1 This is a flowchart of a gesture interaction control method disclosed in this application;
[0035] Figure 2 This is a flowchart of a specific gesture interaction control method disclosed in this application;
[0036] Figure 3 This is a flowchart of another specific gesture interaction control method disclosed in this application;
[0037] Figure 4 This is a gesture diagram disclosed in this application;
[0038] Figure 5 This is a schematic diagram of the structure of a gesture interaction control device disclosed in this application;
[0039] Figure 6 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0041] With the continuous development of software technology, many web-based medical image browsing front-ends have emerged in the cloud imaging industry, such as the open-source ohif. The existing technologies are all based on cornerstone technology, which can display images well, but there are the following problems: when consultations or teaching are needed on a large screen, everyone needs to go to a designated location to operate the cloud film system with a mouse, which is inconvenient and difficult to understand.
[0042] Therefore, this application provides a gesture interaction control scheme that enables free operation of the cloud film system without the aid of external devices such as a mouse.
[0043] Reference Figure 1 As shown, this embodiment of the invention discloses a gesture interaction control method, characterized in that it is applied to a cloud film system and includes:
[0044] Step S11: When the cloud film system detects that gesture recognition operation is enabled, acquire the video stream of the captured hand movements.
[0045] In this embodiment, a camera is pre-set as the target camera for capturing hand movements. This camera may include, but is not limited to, a camera built into the computer device with the cloud film system installed, or an external ordinary camera connected to the cloud film system. The number of cameras can be one or more, without specific limitation. After determining the target camera, its specific parameters are set, including but not limited to, size parameters and position information parameters. When the cloud film system is detected to be in gesture recognition mode, the set target camera is used to capture hand movements to generate a corresponding video stream. Specifically, the cloud film system is first opened and switched to gesture recognition mode. When the current operation mode is detected as gesture recognition mode, the target camera is triggered to start capturing. The target camera captures images of hand movements within the current window's field of view according to the preset parameter settings, forming a video stream.
[0046] Step S12: Perform video frame extraction on the video stream to determine the target video frame.
[0047] In this embodiment, video frame segmentation and image conversion processing are performed on the video stream to determine the target video frame. Specifically, since the video stream contains a large amount of hand movement information, it is necessary to perform video frame segmentation to obtain multiple video frames, and then perform image conversion processing on each individual video frame to obtain the target video frame. The video frame segmentation operation can sample and segment the video stream at time intervals, or it can be performed in a custom manner; there is no specific limitation on this.
[0048] In this embodiment, video frame extraction is performed on the video stream to obtain video frames to be processed; bilateral filtering and mirror flipping are then performed on the video frames to be processed sequentially to obtain preprocessed video frames. It is understood that the purpose of bilateral filtering on the video frames to be processed is to smooth them; then, the smoothed video frames are mirror flipped to obtain mirrored video frames, which are used as the preprocessed video frames.
[0049] The preprocessed video frame is cropped to obtain a rectangular region containing the hand outline, and this rectangular region is used as the target video frame. It can be understood that a rectangular region is cropped from the preprocessed video frame to serve as the target video frame; this rectangular region is the area where gesture recognition is required. During the cropping process, image recognition technology is used to determine the approximate hand outline. Based on the location information of the hand outline, the rectangular region to be cropped is determined. The purpose of cropping the rectangular region is to avoid recognizing irrelevant image content and thus avoid affecting subsequent hand movement recognition.
[0050] Step S13: Determine the corresponding hand gesture tag for the target video frame, and perform interactive control of the cloud film system based on the hand gesture tag and the cloud film system control mapping relationship.
[0051] In this embodiment, the target video frame is tagged to obtain the corresponding hand gesture tags. These hand gesture tags may include, but are not limited to: index finger raised and middle finger bent, index and middle fingers raised simultaneously, three fingers together, three fingers spread out, and drawing a circle with a single finger. The cloud film system is interactively controlled based on the hand gesture tags of the target video frame and the control mapping relationship between them. This cloud film system control mapping relationship can be easily and quickly set by designing the necessary execution nodes on the visual business logic design interface and sequentially associating them. Specifically, based on finger movements and the number of fingers, some operation instructions in the cloud film system correspond to these instructions. This control mapping relationship is pre-set and directly saved to the local database of the cloud film system. Gestures are associated with the cornerstone API for operations such as rotation, switching, zooming in, zooming out, and moving, and can be directly invoked when needed. For example: if the detected hand action label for the target video frame is that the index and middle fingers are raised simultaneously, and the pixel distance between the fingertips of the index and middle fingers is less than 50, then it is considered a mouse click; if the detected hand action label for the target video frame is that the index finger is raised and the middle finger is bent, then it corresponds to calling the mouse movement operation of the Cloud Film system, and moving the image using the system's cornerstoneTools.pan.activate; if the detected hand action label for the target video frame is that all five fingers are raised simultaneously, then it corresponds to calling the Cloud Film system's cornerstoneTools.stackScroll.activate to switch images, and if the five fingers are pointing to the left, it indicates a swipe to the left. The corresponding actions will switch to the previous image. A five-finger gesture pointing to the right indicates swiping right, and swiping right will switch to the next image. If the hand gesture label of the current target video frame is "extending two fingers", the corresponding action will call the cloud film system's cornerstoneTools.pan.activate to move the image. If the hand gesture label of the current target video frame is "three fingers together or outward", the corresponding action will call the cloud film system's cornerstoneTools.zoom.activate to zoom the image. If the hand gesture label of the current target video frame is "a single finger making a circle", the corresponding action will call cornerstone.setViewport to rotate the image at the corresponding angle.
[0052] Therefore, this application discloses a method for acquiring a video stream of hand movements when the cloud film system is activated for gesture recognition; performing video frame segmentation on the video stream to determine a target video frame; determining the corresponding hand movement tag for the target video frame; and interactively controlling the cloud film system based on the mapping relationship between the hand movement tag and the cloud film system control. It is evident that operating the cloud film system through gesture recognition facilitates efficient communication for doctors. Gesture operation allows for more vivid and intuitive manipulation of the cloud film system, replacing traditional mouse operation with user-friendly gestures and offering greater freedom.
[0053] Reference Figure 2 As shown, this embodiment of the invention discloses a specific gesture interaction control method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically:
[0054] Step S21: When the cloud film system detects that gesture recognition operation is enabled, acquire the video stream of the hand movements.
[0055] Step S22: Perform video frame extraction on the video stream, remove background from the extracted video frames, and sequentially perform image grayscale processing, filtering, and binarization on the processed video frames to determine the target video frame.
[0056] In this embodiment, the background information of the intercepted video frame is automatically obtained by using image recognition technology. After the background is determined, the background is removed from the video frame through a background removal algorithm to obtain the foreground. In this project, the foreground is the hand image. The foreground is processed. First, the color image is converted into a grayscale image through image grayscale processing. Specifically, skin color detection is performed, and the video frame images that meet the ranges of 7 < H < 20, 28 < S < 256, and 50 < V < 256 in HSV are screened based on the H, S, V range screening method in the HSV color space. Then, Gaussian filtering is performed to remove noise. Finally, binaryzation processing is performed to obtain a black-and-white image, that is, the target video frame. Among them, the following apis are used in the process of obtaining the target video frame: cv2.createBackgroundSubtractorMOG2(0, bgSubThreshold) to obtain a background model; bgModel.apply(frame, learningrate = ) frame is the newly obtained image. This step is to obtain a foreground mask from the new picture (the foreground is white and the background is black); cv2.erode(fgmask, kernel, iterations = 1) erosion operation, convolution operation, to remove noise; cv2.bitwise_and(frame, frame, mask = fgmask) The mask and the new image are ANDed. White is 1 and black is 0. Any value ANDed with 0 is 0, and any value ANDed with 1 remains unchanged. In this way, the foreground can be cut out. cv2.cvtColor() is used to convert to a grayscale image; cv2.GaussianBlur() is used for Gaussian filtering; cv2.threshold() is used to convert to a binary image.
[0057] Step S23: Determine the corresponding hand action label of the target video frame, and perform interactive control on the cloud film system based on the mapping relationship between the hand action label and the cloud film system control.
[0058] Among them, for the more detailed processing procedures in steps S21 and S23, please refer to the content of the previously disclosed embodiments, and will not be elaborated here.
[0059] Thus, through image processing of the intercepted image, a hand action image without other background information, that is, the target video frame, is obtained, so that the image information in the obtained target video frame is clearer, and it is more convenient to label the hand action and to judge the action of the target video frame.
[0060] Refer to Figure 3 As shown, the embodiment of the present invention discloses a specific gesture interaction control method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically:
[0061] Step S31: When the cloud film system detects that gesture recognition operation is enabled, acquire the video stream of the captured hand movements.
[0062] Step S32: Perform video frame extraction on the video stream to determine the target video frame.
[0063] For more detailed processing steps S31 and S32, please refer to the aforementioned disclosed embodiments; they will not be repeated here.
[0064] Step S33: Obtain the hand indentation of the target video frame, determine the number of finger pits based on the hand indentation; perform a tagging operation on the target video frame based on the number of finger pits to determine the corresponding hand action tag of the target video frame.
[0065] In this embodiment, the label type of the target video frame is determined, that is, the label type of the binary image is determined. The determination method may include, but is not limited to, simple feature-based methods and advanced deep learning methods. This embodiment uses the feature-based method. (Refer to...) Figure 4 As shown in the gesture diagram, each feature of the hand is pre-marked as a different feature point. Then, based on the detection of these feature points, the hand contour and convex hull are obtained using APIs. The convex hull is the convex polygon that the hand perfectly encloses. The hand contour and convex hull can be used to obtain the depressions. By recording the hand depressions, the number of finger pits can be obtained, thus determining the number of fingers. It's important to note that not all hand depressions are finger pits; the angle between the depressions should be less than 90 degrees. Therefore, the cosine theorem is used to calculate the angle of the hand depressions, and depressions with angles less than 90 degrees are identified as finger pits. The APIs used in this process are as follows:
[0066] cv2.findContours() retrieves the contours; note the properties of the returned parameter.
[0067] cv2.convexHull() returns the convex hull, which is essentially a set of points.
[0068] cv2.drawContours() draws the outline.
[0069] cv2.convexityDefects() generates the concavity.
[0070] math.sqrt() calculates the square.
[0071] OpenCV treats an image as a matrix, with each element representing a color.
[0072] Step S34: Track the hand movement trajectory according to the hand action tag of each target video frame, and determine the correspondence between the coordinate point of the finger in the target video frame and the target coordinate point corresponding to the control operation of the cloud film system based on the hand movement trajectory; control the cloud film system based on the correspondence to complete the interactive control of the cloud film system.
[0073] In this embodiment, since the operation of the cloud film system via gesture interaction is performed through hand movements, and in actual application scenarios, a smooth video of hand movements is often used to manipulate the cloud film system, after obtaining the hand movement tag of a single target video frame to control a single command operation of the cloud film system, in order to more intuitively and smoothly demonstrate the gesture interaction, it is necessary to continue to track the hand movement trajectory based on the hand movement tag of each target video frame. Specifically, tracking the hand movement trajectory can be done using image tracking algorithms, which may include, but are not limited to, optical flow, camshift, KCF, deep learning, etc. The trajectory is tracked directly using the hand position detected in each target video frame. If the index finger is detected to be raised and the middle finger is detected to be bent, then it is considered that the mouse is moving. The mouse position coordinates are the position coordinates of the index fingertip, and the hand position is represented by the coordinates of the upper left corner of the bounding rectangle of the hand. Then, the coordinate point sequence is obtained by sampling at equal intervals, and then discretized according to the angle between each coordinate and the center point of the coordinate sequence. Tracking the hand movement trajectory can determine the coordinate points of the fingers in the browser to replace mouse clicks to operate the cloud film system.
[0074] Therefore, by detecting hand indentations in video frame images to obtain the number of hand indentations and determining the number of finger pits based on these indentations, the current target video frame can be tagged. Then, according to the preset control mapping relationship, the relevant operation commands of the cloud film system are called to perform related operations on the images displayed by the cloud film system. Compared with traditional mouse operation, gesture interaction operation eliminates the need for operators to go to a designated location to operate the cloud film system with a mouse, making it more convenient. Furthermore, when operating the cloud film system, corresponding rotation, zoom, and switching operations based on hand movements vividly demonstrate the human body structure in the cloud film, making it easier to understand.
[0075] Reference Figure 5 As shown, this embodiment of the invention also discloses a gesture interaction control device applied to a cloud film system, comprising:
[0076] The video stream acquisition module 11 is used to acquire a video stream of hand movements when the cloud film system detects that the gesture recognition operation is enabled.
[0077] The video frame extraction module 12 is used to perform video frame extraction operations on the video stream to determine the target video frame;
[0078] The interactive control module 13 is used to determine the corresponding hand action tag of the target video frame and to perform interactive control of the cloud film system based on the hand action tag and the cloud film system control mapping relationship.
[0079] In some specific embodiments, the video frame extraction module 12 may specifically include:
[0080] The image conversion submodule is used to perform video frame extraction and image conversion processing on the video stream to determine the target video frame.
[0081] In some specific embodiments, the image conversion submodule may specifically include:
[0082] An image cropping unit is used to perform video frame cropping operations on the video stream to obtain video frames to be processed.
[0083] The video frames to be processed are sequentially subjected to bilateral filtering and mirror flipping to obtain preprocessed video frames.
[0084] In some specific embodiments, the video frame extraction module 12 may specifically include:
[0085] The video frame cropping unit is used to crop the preprocessed video frame to obtain a rectangular area containing the outline of the hand, and to use the rectangular area as the target video frame.
[0086] In some specific embodiments, the video frame extraction module 12 may specifically include:
[0087] The background removal unit is used to perform background removal processing on the rectangular region, and then perform image grayscale processing, filtering processing and binarization processing on the processed rectangular region in sequence to determine the target video frame.
[0088] In some specific embodiments, the interactive control module 13 may specifically include:
[0089] A tag determination unit is used to acquire the hand indentation of the target video frame and determine the number of finger pits based on the hand indentation;
[0090] The target video frame is tagged based on the number of finger sockets to determine the corresponding hand action tag for the target video frame.
[0091] In some specific embodiments, the interactive control module 13 may specifically include:
[0092] An interactive control unit is used to track the hand movement trajectory based on the hand action tag of each target video frame, and determine the correspondence between the coordinate points of the finger in the target video frame and the target coordinate points corresponding to the control operation of the cloud film system based on the hand movement trajectory.
[0093] Based on the aforementioned correspondence, the cloud film system is controlled to complete the interactive control of the cloud film system.
[0094] Therefore, this application discloses a method for acquiring a video stream of hand movements when the cloud film system is activated for gesture recognition; performing video frame segmentation on the video stream to determine a target video frame; determining the corresponding hand movement tag for the target video frame; and interactively controlling the cloud film system based on the mapping relationship between the hand movement tag and the cloud film system control. It is evident that operating the cloud film system through gesture recognition facilitates efficient communication for doctors. Gesture operation allows for more vivid and intuitive manipulation of the cloud film system, replacing traditional mouse operation with user-friendly gestures and offering greater freedom.
[0095] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0096] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the gesture interaction control method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0097] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0098] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0099] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0100] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the gesture interaction control method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.
[0101] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed gesture interaction control method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0102] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0103] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described functions using different methods for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. Software modules can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.
[0104] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0105] The foregoing has provided a detailed description of the gesture interaction control method, apparatus, device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A gesture interaction control method, characterized in that, Applications in cloud film systems include: When the cloud film system detects that gesture recognition is enabled, a video stream of hand movements is captured. Perform video frame extraction on the video stream to determine the target video frame; The corresponding hand gesture tags for the target video frame are determined, and the cloud film system is interactively controlled based on the hand gesture tags and the control mapping relationship between the cloud film system and the cloud film system. The hand gesture tags include one or more of the following: index finger raised and middle finger bent, index finger and middle finger raised at the same time, three fingers together, three fingers spread out, one finger drawing a circle, five fingers raised at the same time, and two fingers extended. The cloud film system control mapping relationship includes: if the hand action label is that the index finger and middle finger are raised simultaneously, and the pixel distance between the fingertips of the index finger and the middle finger is less than 50, it corresponds to a mouse click operation; if the hand action label is that the index finger is raised and the middle finger is bent, it corresponds to a mouse movement operation; if the hand action label is that all five fingers are raised simultaneously, it corresponds to an image switching operation; if the hand action label is that two fingers are extended, it corresponds to an image movement operation; if the hand action label is that three fingers are together or three fingers are spread out, it corresponds to an image zoom operation; if the hand action label is that one finger is drawing a circle, it corresponds to an image selection operation.
2. The gesture interaction control method according to claim 1, characterized in that, The step of performing video frame extraction on the video stream to determine the target video frame includes: The video stream is subjected to video frame extraction and image conversion processing to determine the target video frame.
3. The gesture interaction control method according to claim 2, characterized in that, The step of performing video frame extraction and image conversion processing on the video stream includes: Perform video frame extraction on the video stream to obtain the video frames to be processed; The video frames to be processed are sequentially subjected to bilateral filtering and mirror flipping to obtain preprocessed video frames.
4. The gesture interaction control method according to claim 3, characterized in that, The determination of the target video frame includes: The preprocessed video frame is cropped to obtain a rectangular region containing the outline of the hand, and the rectangular region is used as the target video frame.
5. The gesture interaction control method according to claim 4, characterized in that, The determination of the target video frame includes: The rectangular region is subjected to background removal processing, and the processed rectangular region is then subjected to image grayscale processing, filtering processing, and binarization processing in sequence to determine the target video frame.
6. The gesture interaction control method according to any one of claims 1 to 5, characterized in that, Determining the corresponding hand gesture tag for the target video frame includes: Obtain the hand indentation of the target video frame, and determine the number of finger pits based on the hand indentation; The target video frame is tagged based on the number of finger sockets to determine the corresponding hand action tag for the target video frame.
7. The gesture interaction control method according to claim 1, characterized in that, The interactive control of the cloud film system based on the mapping relationship between the hand gesture tags and the cloud film system control includes: The hand movement trajectory is tracked according to the hand action tag of each target video frame, and the correspondence between the coordinate points of the finger in the target video frame and the target coordinate points corresponding to the control operation of the cloud film system is determined based on the hand movement trajectory. Based on the aforementioned correspondence, the cloud film system is controlled to complete the interactive control of the cloud film system.
8. A gesture interaction control device, characterized in that, Applications in cloud film systems include: The video stream acquisition module is used to acquire a video stream of hand movements when the cloud film system detects that gesture recognition operation has been enabled. The video frame extraction module is used to perform video frame extraction operations on the video stream to determine the target video frame; An interactive control module is used to determine the corresponding hand gesture label of the target video frame and to interactively control the cloud film system based on the hand gesture label and the control mapping relationship between the cloud film system; the hand gesture label includes one or more of the following: index finger raised and middle finger bent, index finger and middle finger raised at the same time, three fingers together, three fingers spread out, one finger drawing a circle, five fingers raised at the same time, and two fingers extended. The cloud film system control mapping relationship includes: if the hand action label is that the index finger and middle finger are raised simultaneously, and the pixel distance between the fingertips of the index finger and the middle finger is less than 50, it corresponds to a mouse click operation; if the hand action label is that the index finger is raised and the middle finger is bent, it corresponds to a mouse movement operation; if the hand action label is that all five fingers are raised simultaneously, it corresponds to an image switching operation; if the hand action label is that two fingers are extended, it corresponds to an image movement operation; if the hand action label is that three fingers are together or three fingers are spread out, it corresponds to an image zoom operation; if the hand action label is that one finger is drawing a circle, it corresponds to an image selection operation.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the gesture interaction control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the gesture interaction control method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Medical image browse device with somatosensory interaction mode
CN102354345A
Visual interactive interface control method, system and device and storage medium
CN114217728A