Video behavior recognition and modification
A multi-granularity video action recognition process using a content-aware and multi-scale skeletal 3D graph convolutional network addresses the challenge of varying action scales, improving recognition accuracy by categorizing and linking skeletal points for precise video action representation.
Patent Information
- Application Number
- JP2023573345
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-23
- Filing Date
- 2022-05-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-05-11
AI Technical Summary
Existing video action recognition technologies face challenges in accurately representing actions of different levels of granularity due to the use of single-scale skeletons, leading to difficulties in precise video action representation and recognition.
A multi-granularity video action recognition process is implemented using a content-aware and multi-scale skeletal 3D graph convolutional network, which extracts and categorizes skeletal points at multiple digital levels, adjusts visual window sizes based on movement distance, and links feature vectors to point coordinates for improved recognition accuracy.
The solution enables accurate video action recognition by adapting to different action granularities, enhancing the precision and effectiveness of video action representation.
Smart Images

Figure 0007720926000025 
Figure 0007720926000026 
Figure 0007720926000027
Abstract
Description
[Background technology]
[0001] The present invention generally relates to methods for automating video action recognition and modification, and more particularly to methods and related systems for improving video and software technologies related to extracting and categorizing skeleton points from a video stream associated with a video representation of a user performing a user movement action, generating an initial visual window encompassing a group of skeleton points, extracting feature vectors and linking them to point coordinates, and enabling the video stream for accurate presentation-related video action recognition. A typical skeleton-based graph convolutional network is associated with video action recognition in which a skeleton is first extracted based on the OpenPose process. Similarly, a spatial information-based graph convolutional network uses position and confidence attributes to extract temporal information based on a temporal convolutional network. The process of using position and confidence attributes may tend to bypass information related to the skeleton points. Furthermore, a typical video stream may be associated with actions of different levels of granularity, which may make it difficult to represent the actions based on a one-scale skeleton. Accordingly, the present method and related system are configured to enable a multi-granularity video action recognition process based on a content-aware and multi-scale skeletal 3D graph convolutional network for improving recognition accuracy associated with multi-granularity video recognition actions. Summary of the Invention
[0002] A first aspect of the present invention is a hardware device including a processor coupled to a computer readable memory unit, the memory unit including instructions that, when executed by the processor, implement a video motion recognition and modification method, the method including: receiving, by the processor, a video stream including user motion movements; extracting, by the processor, from the video stream via a plurality of sensors, skeletal points associated with a video representation of a user performing the user motion movements; categorizing, by the processor, the skeletal points with respect to a number of digital levels; generating, by the processor, an initial visual window surrounding a group of skeletal points in a plurality of video frames of the video stream; determining, by the processor, an average movement distance of the group of skeletal points with respect to the plurality of video frames; adjusting, by the processor, a size of the visual window with respect to each of the plurality of video frames in response to a result of the determination; and a processor responsive to a result of enabling the convolutional neural network, enabling the video stream for video action recognition associated with accurate presentation of the video stream.
[0003] Some embodiments of the present invention further provide a hardware device for extracting surrounding features associated with the skeleton points. Similarly, some embodiments of the present invention are configured to categorize and link the skeleton points with respect to multiple digital features to determine classification probabilities associated with the operation of a convolutional neural network. Advantageously, these embodiments provide an effective means for accurately enabling a multi-granularity video action recognition process based on a content-aware and multi-scale skeletal 3D graph convolutional network for improving recognition accuracy associated with multi-granularity video recognition actions.
[0004] A second aspect of the present invention is a method for video motion recognition and modification, comprising: receiving, by a processor of a hardware device, a video stream including user motion movements; extracting, by the processor, from the video stream via a plurality of sensors, skeletal points associated with a video representation of a user performing the user motion movements; categorizing, by the processor, the skeletal points with respect to a number of digital levels; generating, by the processor, an initial visual window encompassing a group of skeletal points in a plurality of video frames of the video stream; determining, by the processor, an average distance traveled by the group of skeletal points with respect to the plurality of video frames; and, responsive to the results of this determination, generating, by the processor, an initial visual window encompassing a group of skeletal points with respect to each of the plurality of video frames. a processor executing Scale Invariant Feature Transform (SIFT) code to adjust a size of a set of skeletal points in response to a result of the adjustment; extracting, by a processor executing OpenPose code, feature vectors of a set of skeletal points; extracting, by a processor executing OpenPose code, point coordinates of the skeletal points; performing, by the processor, a first linking of the feature vectors to the point coordinates; generating, by the processor, a convolutional neural network associated with the linking of the feature vectors to the point coordinates; and enabling, by the processor, in response to a result of enabling the convolutional neural network, the video stream for video action recognition associated with accurate presentation of the video stream.
[0005] Some embodiments of the present invention further provide methods for extracting surrounding features associated with skeleton points. Similarly, some embodiments of the present invention are configured to categorize and link skeleton points with respect to multiple digital features to determine classification probabilities associated with the operation of a convolutional neural network. Advantageously, these embodiments provide an effective means for accurately enabling a multi-granularity video action recognition process based on content-aware and multi-scale skeletal 3D graph convolutional networks to improve recognition accuracy associated with multi-granularity video recognition actions.
[0006] A third aspect of the present invention is a computer program product comprising a computer readable hardware storage device having computer readable program code stored thereon, the computer readable program code comprising an algorithm which, when executed by a processor of the hardware device, implements a video movement recognition and modification method, the method comprising: receiving, by the processor, a video stream comprising user movement movements; extracting, by the processor, from the video stream via a plurality of sensors, skeletal points associated with a video representation of a user performing the user movement movements; categorizing, by the processor, the skeletal points with respect to a number of digital levels; generating, by the processor, an initial visual window surrounding a group of skeletal points in a plurality of video frames of the video stream; determining an average distance of movement of skeletal points, adjusting, by a processor responsive to a result of the determination, a size of a visual window for each of a plurality of video frames, extracting, by a processor running Scale Invariant Feature Transform (SIFT) code, feature vectors of a group of skeletal points in response to a result of the adjustment, extracting, by a processor running OpenPose code, point coordinates of the skeletal points, performing, by the processor, a first linking of the feature vectors to the point coordinates, generating, by the processor, a convolutional neural network associated with the linking of the feature vectors to the point coordinates, and enabling, by the processor responsive to a result of enabling the convolutional neural network, the video stream for video action recognition associated with accurate presentation of the video stream.
[0007] Some embodiments of the present invention further provide a computer program product for extracting surrounding features associated with skeleton points. Similarly, some embodiments of the present invention are configured to categorize and link skeleton points with respect to multiple digital features to determine classification probabilities associated with the operation of a convolutional neural network. Advantageously, these embodiments provide an effective means for accurately enabling a multi-granularity video action recognition process based on a content-aware and multi-scale skeletal 3D graph convolutional network for improving recognition accuracy associated with multi-granularity video recognition actions.
[0008] Advantageously, the present invention provides a simple method and associated system that can automate the recognition and modification of video actions. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates a system according to an embodiment of the present invention for improving video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motor action, generating an initial visual window surrounding a group of skeletal points, extracting and linking feature vectors to point coordinates, and enabling the video stream for video action recognition related to accurate presentation. [Figure 2] FIG. 2 illustrates an algorithm detailing a process flow according to an embodiment of the present invention enabled by the system of FIG. 1 for improving video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motor action, generating an initial visual window surrounding a group of skeletal points, extracting and linking feature vectors to point coordinates, and enabling the video stream for video action recognition related to accurate presentation. [Figure 3] 2 is a diagram showing the internal structure of the software / hardware of FIG. 1 according to an embodiment of the present invention. [Figure 4] 2 illustrates an initial visual window generation process enabled by the system of FIG. 1 in accordance with an embodiment of the present invention. [Figure 5] FIG. 2 illustrates a point coordinate extraction process enabled by the system of FIG. 1 in accordance with an embodiment of the present invention. [Figure 6] FIG. 2 illustrates a computer system according to an embodiment of the present invention that may be used by the system of FIG. 1 to refine video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user movement action, generating an initial visual window encompassing a group of skeletal points, extracting and linking feature vectors to point coordinates, and enabling the video stream for video action recognition related to accurate presentation. [Figure 7] FIG. 1 illustrates a cloud computing environment in accordance with an embodiment of the present invention. [Figure 8] FIG. 1 illustrates a set of functional abstraction layers provided by a cloud computing environment in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] FIG. 1 illustrates a system 100 according to an embodiment of the present invention for improving video and software technologies related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motion action, generating an initial visual window encompassing a group of skeletal points, extracting feature vectors and linking them to point coordinates, and enabling the video stream for video action recognition associated with accurate presentation. Exemplary skeleton-based graph convolutional networks are associated with video action recognition in which a skeleton is first extracted. Similarly, exemplary spatial information-based graph convolutional networks are associated with temporal convolutional networks, where temporal information is extracted using position and confidence attributes. Furthermore, the process of using position and confidence attributes may tend to bypass information associated with skeletal points. Furthermore, typical video streams may be associated with actions at different levels of granularity, making it difficult to represent actions based on a single-scale skeleton. Thus, the system 100 is configured to recognize multi-granularity video actions, perform an adaptive surrounding feature extraction process based on skeleton points, and enable a multi-scale skeleton 3D graph convolutional network, thereby improving the recognition accuracy for multi-granularity actions and providing an accurate method for extracting video recognition attributes.
[0011] System 100 enables retrieval of information associated with skeleton points for image information in such a manner that skeleton points associated with different scales can be configured to represent different granularity video actions, which can be associated with a skeleton-based three-dimensional graph.
[0012] The system 100 is configured to perform a multi-granularity video action recognition process based on content awareness of a multi-scale skeletal 3D graph based on a convolutional network. This process begins when a skeleton is extracted based on the OpenPose system (i.e., a real-time multi-person system that jointly detects human body, hand, face, and foot keypoints (e.g., a total of 135 keypoints) for a single image). Similarly, the system 100 includes a feature extraction component associated with the skeleton points to construct a 3D graph based on skeletal attributes at different scales.
[0013] The system 100 of FIG. 1 includes hardware device 139, video hardware 114, database 115, and network interface controller 153 interconnected via network 117. Hardware device 139 includes sensors 112, circuitry / logic 127, and software / hardware 121. Video hardware 114 may include a remote video source system (e.g., a video storage system, a video streaming system, a video projector, etc.). Hardware device 139 and video hardware 114 may each include an embedded device. An embedded device is defined herein as a special-purpose device or computer that includes a combination of computer hardware and software (fixed-function or programmable) specifically designed to perform a specific application function. A programmable embedded computer or device may include an application-specific programming interface. In one embodiment, hardware device 139 and video hardware 114 may each comprise application-specific hardware devices including application-specific (non-general-purpose) hardware and circuitry (i.e., application-specific, discrete, non-general-purpose analog, digital, and logic-based circuitry) for performing (independently or in combination) the processes described with respect to Figures 1-8. Those application-specific, discrete, non-general-purpose analog, digital, and logic-based circuitry (e.g., sensors 112, circuitry / logic 127, software / hardware 121, etc.) may include specially designed proprietary components (e.g., application-specific integrated circuits, such as application-specific integrated circuits (ASICs) designed solely to perform automated processes for improving video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motor action, generating an initial visual window surrounding a group of skeletal points, extracting feature vectors and linking them to point coordinates, and enabling the video stream for video action recognition related to accurate presentation).The sensors 112 may include any type of internal or external sensor, including, among other things, GPS sensors, Bluetooth beaconing sensors, cellular phone detection sensors, Wi-Fi positioning detection sensors, triangulation detection sensors, activity tracking sensors, temperature sensors, ultrasonic sensors, light sensors, video search devices, humidity sensors, voltage sensors, network traffic sensors, etc. The network 117 may include any type of network, including, among other things, a local area network (LAN), a wide area network (WAN), the Internet, a wireless network, etc.
[0014] The system 100 is enabled to perform a process for extracting visual ambient features by enabling adaptive ambient feature extraction code for the Scale Invariant Feature Transform (SIFT), which includes a computer vision related feature detection algorithm for detecting and describing local features within an image. This process includes: 1. Changing the initial window size 5 to extract video features. 2. Set the frame strap size to 10 and determine the average movement distance of each skeleton point. 3. Using the relative distance to adjust the visual window size. 4. Extracting box features for 1D vectors using a SIFT related process.
[0015] The system 100 is further enabled to execute a process for implementing a multi-scale skeletal 3D graph convolutional network modeling process, which is performed as follows: 1. Undersample three skeletal layers based on the OpenPose extracted skeleton. 2. Generate a 3D skeletal graph. 3. Run a 3D graph convolutional network to extract video features at each layer. 4. Fusing 3D graph convolutional network features with attention models for different layers.
[0016] FIG. 2 illustrates an algorithm detailing a process flow according to an embodiment of the present invention, enabled by the system 100 of FIG. 1 for improving video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motion, generating an initial visual window encompassing the skeletal points, extracting feature vectors and linking them to point coordinates, and enabling the video stream for accurate presentation-related video action recognition. Each of the steps of the algorithm of FIG. 2 can be enabled and performed in any order by a computer processor executing computer code. Furthermore, each of the steps of the algorithm of FIG. 2 can be enabled and performed in combination by the hardware device 139 and the video hardware 114. In step 200, a video stream including a user motion is received by the hardware device. In step 202, skeletal points are extracted from the video stream via a sensor. The skeletal points are associated with the video representation of the user performing the user motion. In step 204, the skeletal points are categorized with respect to multiple digital levels. The skeleton points can be upsampled to a number of digital levels.
[0017] In step 208, an initial visual window is generated within a video frame of the video stream that surrounds a group of skeleton points. In step 210, an average movement distance of the skeleton points is determined with respect to the video frame. Similarly, surrounding features associated with the skeleton points can be extracted in such a manner that the determination of the average movement distance of the skeleton points is further based on the surrounding features. The determination of the average movement distance is calculated using the formula
number
number
number
number
[0018] In step 212, the size of the visual window is adjusted for each of the video frames (based on the results of step 210). The adjustment of the visual window size is performed using the formula
number
number
number
number
[0019] In step 214, feature vectors of the skeleton points are extracted by executing Scale Invariant Feature Transform (SIFT) code in response to the results of step 212. In step 216, point coordinates of the skeleton points are extracted by executing OpenPose code. In step 218, the feature vectors are linked to the point coordinates. In step 220, a convolutional (digital) neural network is generated. The convolutional (digital) neural network is associated with linking the feature vectors to the point coordinates. In step 224, the skeleton points are categorized and linked with respect to multiple digital features, and classification probabilities associated with the operation of the convolutional neural network are determined in response to the results of the linking step. Similarly, the skeleton points can be categorized with respect to multiple digital features, and the multiple digital features can be linked. Furthermore, classification probabilities associated with the operation of the convolutional neural network can be determined based on the linking step.
[0020] In step 228, the video stream is enabled for video action recognition related to accurate presentation of the video stream.
[0021] FIG. 3 shows an internal structural diagram of the software / hardware 121 of FIG. 1 in accordance with an embodiment of the present invention. The software / hardware 121 includes an extraction and categorization module 304, an adjustment module 310, a linking module 308, a video enablement module 314, and a communications controller 302. The extraction and categorization module 304 includes specialized hardware and software for controlling all functions related to the extraction and categorization step of FIG. 2. The adjustment module 310 includes specialized hardware and software for controlling all functions related to the adjustment step described with respect to the algorithm of FIG. 2. The linking module 308 includes specialized hardware and software for controlling all functions related to the linking step of FIG. 2. The video enablement module 314 includes specialized hardware and software for controlling all functions related to the video stream enablement and presentation step of the algorithm of FIG. 2. The communications controller 302 is enabled to control all communications between the extraction and categorization module 304, the adjustment module 310, the linking module 308, and the video enablement module 314.
[0022] FIG. 4 illustrates an initial visual window generation process 400 enabled by the system 100 of FIG. 1, according to an embodiment of the present invention. Frame 401 includes a first sample frame containing skeleton points 402a...402n and connections 403a...403n connecting the skeleton points 402a...402n. The OpenPose tool is executed to extract the skeleton points 402a...402n for every frame. Frame 404 includes a second sample frame containing an original window 404a surrounding a group 404b of skeleton points 402a...402n, associated with extracting contextual features to represent the skeleton points 402a...402n. Frames 408a...408n include frames associated with a process for determining reshape factors for the width and height of the original window. This process enables the use of ten frames surrounding the current frame to calculate the Euclidean distance of a current skeleton point (among the skeleton points 402a...402n) located between the two concatenated frames. Similarly, the average Euclidean distance is determined among the 10 frames for the reshape factor (as shown in relation to the formula described with reference to FIG. 2). Frame 410 includes sample frames associated with the process for reshaping the original window using the reshape factor described above (as shown in relation to the formula described with reference to FIG. 2). Frames 412 and 414 include sample frames associated with the process for extracting context features using the SIFT algorithm for skeleton points within window 404a. This process for extracting context features is performed to generate feature vectors for the skeleton points within window 404a. These feature vectors are used as input data for the deep learning model.
[0023] FIG. 5 illustrates a point coordinate extraction process 500 enabled by the system 100 of FIG. 1 in accordance with an embodiment of the present invention. For each frame 502, skeleton points are extracted by executing the OpenPose tool. The OpenPose tool is configured to extract 18 skeleton points for one person in frame 502. Block 504 illustrates the skeleton points being downsampled to different levels L1, L2, and L3 to extract more information from each frame 502. For example, at level L2, the midpoint between every two skeleton points is assigned as a skeleton point, thereby enabling a process to extract SIFT features around the new skeleton points. Block 506 illustrates the operation of a spatio-temporal graph convolutional network (ST-GCN) algorithm, which is implemented by executing code and applying a deep neural network backbone to a graph (vertex and edge) sequence instead of an image. Block 510 illustrates a process for feeding the extracted SIFT features into an ST-GCN model such that one feature vector is extracted per skeleton point level. Therefore, every level is configured to concatenate multiple vectors into one large vector. Similarly, a linear layer is implemented for the neural network to reduce the vector dimension to the same size as the number of action classes. Block 508 shows a SoftMax layer for mapping the float vectors to class probabilities in such a way that the maximum probability in the final vector contains the recognized action.
[0024] FIG. 6 illustrates a computer system 90 (e.g., hardware device 139 and video hardware 114 of FIG. 1 ) according to an embodiment of the present invention that may be used by or include system 100 of FIG. 1 to improve video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motor action, generating an initial visual window surrounding a group of skeletal points, extracting and linking feature vectors to point coordinates, and enabling the video stream for video action recognition related to accurate presentation.
[0025] Aspects of the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to herein as a "circuit," "module," or "system."
[0026] The present invention may be a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0027] The computer-readable storage medium may be any tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination thereof. As used herein, computer-readable storage media should not be construed as being themselves ephemeral signals, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating in waveguides or other transmission media (e.g., light pulses traveling in fiber optic cables), or electrical signals transmitted over electrical wires.
[0028] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each corresponding computing / processing device, or can be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. This network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium in each corresponding computing / processing device for storage.
[0029] The computer-readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, C++, Spark, R, and the like, and traditional procedural programming languages such as the “C” programming language or the like. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or remote server. In the last scenario above, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to carry out aspects of the present invention.
[0030] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0031] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine, where the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may be stored on a computer-readable storage medium and capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0032] The computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0033] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions that implement the specified logical function(s). In some alternative implementations, the functions shown in the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may actually be performed as a single step, or may be performed simultaneously, substantially simultaneously, or with partial or complete overlap in time, or the blocks may sometimes be performed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or executes a combination of dedicated hardware and computer instructions.
[0034] The computer system 90 shown in FIG. 6 includes a processor 91, an input device 92 coupled to the processor 91, an output device 93 coupled to the processor 91, and memory devices 94 and 95, each coupled to the processor 91. The input device 92 can be, among other devices, a keyboard, a mouse, a camera, a touchscreen, etc. The output device 93 can be, among other devices, a printer, a plotter, a computer screen, a magnetic tape, a removable hard disk, a floppy disk, etc. The memory devices 94 and 95 can be, among other devices, a hard disk, a floppy disk, a magnetic tape, optical storage such as a compact disk (CD) or a digital video disk (DVD), a dynamic random access memory (DRAM), a read-only memory (ROM), etc. The memory device 95 contains computer code 97. Computer code 97 includes algorithms (e.g., the algorithm of FIG. 2) for improving video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motor action, generating an initial visual window encompassing a group of skeletal points, extracting and linking feature vectors to point coordinates, and enabling the video stream for video action recognition associated with accurate presentation. Processor 91 executes computer code 97. Memory device 94 includes input data 96. Input data 96 includes inputs required by computer code 97. Output device 93 displays output from computer code 97.One or both memory devices 94 and 95 (or one or more additional memory devices, such as read-only memory device 96) may contain algorithms (e.g., the algorithms of FIG. 2) and may be used as a computer-usable medium (or computer-readable medium or program storage device) having computer-readable program code embodied therein or other data stored thereon, the computer-readable program code including computer code 97. Generally, a computer program product (or article of manufacture) of computer system 90 may include a computer-usable medium (or program storage device).
[0035] In some embodiments, rather than storing and accessing stored computer program code 84 (e.g., including algorithms) on a hard drive, optical disk, or other writable, rewritable, or removable hardware memory device 95, computer program code 84 may be stored on a non-removable, static, read-only storage medium such as a read-only memory (ROM) device 85, or processor 91 may access computer program code 84 directly from such non-removable, static, read-only medium 85. Similarly, in some embodiments, stored computer program code 97 may be stored as computer-readable firmware 85, or processor 91 may access computer program code 97 directly from such firmware 85, rather than from a more dynamic or removable hardware data storage device 95 such as a hard drive or optical disk.
[0036] Additionally, any of the components of the present invention may be generated, integrated, hosted, maintained, deployed, managed, serviced, etc., by a service supplier offering to extract and categorize skeletal points from a video stream associated with a video representation of a user performing user motor actions, generate an initial visual window surrounding a group of skeletal points, extract feature vectors and link them to point coordinates, and refine video and software techniques related to enabling the video stream for accurate presentation-related video action recognition. Accordingly, the present invention discloses a process for deploying, generating, integrating, hosting, maintaining, or integrating a computing infrastructure, or a combination thereof, comprising integrating into a computer system 90 computer-readable code that, in combination with the computer system 90, can perform a method for enabling a process for extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing user motor actions, generate an initial visual window surrounding a group of skeletal points, extract feature vectors and link them to point coordinates, and refine video and software techniques related to enabling the video stream for accurate presentation-related video action recognition. In another embodiment, the present invention provides a commercial method for performing the process steps of the present invention on a subscription, advertising, or fee basis, or a combination thereof, i.e., a service supplier, such as a solutions integrator, can offer to enable a process for improving video and software techniques related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user movement action, generating an initial visual window surrounding a group of skeletal points, extracting feature vectors and linking them to point coordinates, and enabling the video stream for video action recognition related to accurate presentation.In this case, the service supplier may create, maintain, support, etc., the computer infrastructure that performs the process steps of the present invention for one or more customers, and in return, the service supplier may receive compensation from the customer based on a subscription and / or fee agreement and / or may receive compensation by selling advertising content to one or more third parties.
[0037] Although Figure 6 illustrates computer system 90 as a particular configuration of hardware and software, any configuration of hardware and software known to one of ordinary skill in the art can be utilized for the purposes described above in connection with the particular computer system 90 of Figure 6. For example, memory devices 94 and 95 can be part of a single memory device rather than being separate memory devices.
[0038] Cloud Computing Environment Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the invention may be implemented in connection with any other type of computing environment now known or later developed.
[0039] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the provider of the service. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0040] The features are as follows:
[0041] On-Demand Self-Service: Cloud consumers can automatically and unilaterally provision computing capacity, such as server time and network storage, as needed without the need for human interaction with the provider of this service.
[0042] Broad network access: The functionality is available over the network and is accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0043] Resource Pooling: To serve a large number of consumers using a multi-tenant model, the provider's computing resources are pooled, with different physical and virtual resources dynamically allocated and reallocated according to demand. Consumers generally have no control over or knowledge of the exact location of the resources provided to them, but there is a sense of location independence in the sense that they can specify location at a higher level of abstraction (e.g., country, state, or data center).
[0044] Rapid Elasticity: Features can be provisioned rapidly and elastically, sometimes automatically, to scale out quickly, and can be released quickly to scale in quickly. To the consumer, the features available for provisioning often appear infinite, and they can purchase as much or as little as they want, at any time.
[0045] Metered Services: Cloud systems automatically control and optimize resource usage by intervening in metering functions at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and user accounts in use). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services being used.
[0046] The service model is as follows:
[0047] Software as a Service (SaaS): This functionality is offered to consumers, who use a provider's applications running on a cloud infrastructure. These applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). Aside from possibly setting limited user-specific application configuration, the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application functions.
[0048] Platform as a Service (PaaS): This capability is offered to consumers to deploy consumer-created or acquired applications written using provider-supported programming languages and tools on cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the application hosting environment configuration.
[0049] Infrastructure as a Service (IaaS): This functionality is provided to consumers that supplies processing, storage, network, and other basic computing resources on which they can deploy and run any software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating systems, storage, and deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).
[0050] The deployment model is as follows:
[0051] Private Cloud: This cloud infrastructure is operated solely for the organization. The infrastructure can be managed by the organization or a third party and can exist on-premise or off-premise.
[0052] Community Cloud: This cloud infrastructure is shared by several organizations to support a specific community with shared interests (e.g., mission, security requirements, policy and compliance concerns). The infrastructure can be managed by the organizations or a third party and can exist on-premises or off-premises.
[0053] Public Cloud: This cloud infrastructure is available to the general public or large industry groups and is owned by an organization that sells cloud services.
[0054] Hybrid cloud: This cloud infrastructure is a composite of two or more clouds (private, community, or public) that maintain their own distinct entities but are joined together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0055] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0056] Referring now to FIG. 7, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10, with which local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or combinations thereof, can communicate. The nodes 10 can communicate with each other. The nodes may be physically or virtually grouped into one or more networks (not shown), such as the private, community, public, or hybrid clouds described above, or combinations thereof. This enables the cloud computing environment 50 to provide infrastructure, platform, or software, or combinations thereof, as a service, thereby eliminating the need for cloud consumers to maintain resources on their local computing devices. It is understood that the types of computing devices 54A, 54B, 54C and 54N shown in FIG. 7 are intended to be examples only, and that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device over any type of network or addressable network connection or both (e.g., using a web browser).
[0057] Referring now to Figure 8, a set of functional abstraction layers provided by cloud computing environment 50 (see Figure 7) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 8 are intended to be examples only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0058] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0059] The virtualization layer 70 provides an abstraction layer that can provide the following examples of virtual entities: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.
[0060] In one example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification of cloud consumers and tasks and protection of data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 87 provides allocation and management of cloud computing resources such that required service levels are achieved. Service level agreement (SLA) planning and fulfillment 88 provides proactive coordination and procurement of anticipated future cloud computing resource needs in accordance with SLAs.
[0061] The workload layer 101 provides examples of functionality that can utilize a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include mapping and navigation 102, software development and lifecycle management 103, virtual classroom instructional delivery 133, data analysis processing 134, transaction processing 106, and video and software refinement techniques 107 related to extracting and categorizing skeletal points from a video stream associated with a video representation of a user performing a user motion action, generating an initial visual window surrounding a group of skeletal points, extracting feature vectors and linking them to point coordinates, and enabling the video stream for video action recognition related to accurate presentation.
[0062] While embodiments of the present invention have been described herein for purposes of illustration, many changes and modifications will become apparent to those skilled in the art. It is therefore intended in the appended claims to cover all such changes and modifications that fall within the scope of this invention.
Claims
1. 1. A hardware device including a processor coupled to a computer-readable memory unit, the memory unit containing instructions that, when executed by the processor, implement a video motion recognition and modification method, the method comprising: receiving, by said processor, a video stream including user athletic movements; extracting, by the processor, skeleton points associated with a video representation of a user performing the user motion movements from the video stream via a plurality of sensors; categorizing, by said processor, said skeleton points with respect to a number of digital levels; generating, by the processor, an initial visual window surrounding a group of the skeleton points in a plurality of video frames of the video stream; determining, by the processor, an average movement distance of the set of skeleton points with respect to the plurality of video frames; adjusting, by the processor responsive to a result of said determining, a size of said visual window for each of said plurality of video frames; extracting, by the processor running scale invariant feature transform (SIFT) code in response to the results of the adjustment, a feature vector for the set of skeleton points; extracting, by the processor running OpenPose code, point coordinates of the skeleton points; performing, by the processor, a first linking of the feature vectors to the point coordinates; generating, by the processor, a convolutional neural network associated with said linking, linking said feature vectors to said point coordinates; and enabling, by the processor responsive to a result of enabling the convolutional neural network, the video stream for video action recognition related to accurate presentation of the video stream. Hardware devices, including
2. The method further comprises: extracting, by the processor, surrounding features associated with the skeleton points, wherein the determination of the average movement distance of the group of skeleton points is further based on the surrounding features.
10. The hardware device of claim 1, comprising:
3. The method further comprises: categorizing, by the processor, the skeleton points with respect to a number of digital features; and performing, by said processor, a second linking of said plurality of digital features; 10. The hardware device of claim 1, comprising:
4. The method further comprises: determining, by the processor responsive to results of the first linking and the second linking, classification probabilities associated with operation of the convolutional neural network.
4. The hardware device of claim 3, comprising:
5. said determining said average distance traveled comprising: [Equation 1] performing [Equation 2] where x contains the reshape factor for the original window width (W) of the i-th frame, N contains the number of frames for calculating the reshape factor, dist contains the Euclidean distance, and x t 2. The hardware device of claim 1, wherein x comprises the x coordinate of a skeleton point in frame t.
6. The determination of the average travel distance further comprises: [Equation 3] performing [Equation 4] where yi contains the reshape factor for the original window height (H) of the i-th frame, N contains the number of frames for calculating the reshape factor, dist contains the Euclidean distance, and yi t 6. The hardware device of claim 5, wherein {overscore (x)} comprises the y coordinates of skeleton points within said frame t.
7. said adjusting the size of said visual window comprises: [Equation 5] where W′ comprises the reshaped window width and W comprises the original window width; [Equation 6] 2. The hardware device of claim 1, wherein: ∑ i = 1 i ... comprises a reshape factor for the original window width W of the i th frame.
8. The adjusting of the size of the visual window further comprises: [Equation 7] where H' comprises the reshaped window height and H comprises the original window height; [Equation 8] 8. The hardware device of claim 7, wherein {overscore (x)} comprises a reshape factor for the original window height H of the i-th frame.
9. The method further comprises: upsampling, by said processor, said skeleton points to said multiple digital levels; 10. The hardware device of claim 1, comprising:
10. 1. A method for recognizing and modifying video behavior, comprising: receiving, by a processor of the hardware device, a video stream including user athletic actions; extracting, by the processor, skeleton points associated with a video representation of a user performing the user motion movements from the video stream via a plurality of sensors; categorizing, by said processor, said skeleton points with respect to a number of digital levels; generating, by the processor, an initial visual window surrounding a group of the skeleton points in a plurality of video frames of the video stream; determining, by the processor, an average movement distance of the set of skeleton points with respect to the plurality of video frames; adjusting, by the processor responsive to a result of said determining, a size of said visual window for each of said plurality of video frames; extracting, by the processor running scale invariant feature transform (SIFT) code in response to the results of the adjustment, a feature vector for the set of skeleton points; extracting, by the processor running OpenPose code, point coordinates of the skeleton points; performing, by the processor, a first linking of the feature vectors to the point coordinates; generating, by the processor, a convolutional neural network associated with said linking, linking said feature vectors to said point coordinates; and enabling, by the processor responsive to a result of enabling the convolutional neural network, the video stream for video action recognition related to accurate presentation of the video stream. A method for recognizing and modifying video behavior, including:
11. extracting, by the processor, surrounding features associated with the skeleton points, wherein the determination of the average movement distance of the group of skeleton points is further based on the surrounding features. The method of claim 10 further comprising:
12. categorizing, by the processor, the skeleton points with respect to a number of digital features; and performing, by said processor, a second linking of said plurality of digital features; The method of claim 10 further comprising:
13. determining, by the processor responsive to results of the first linking and the second linking, classification probabilities associated with operation of the convolutional neural network. The method of claim 12 further comprising:
14. said determining said average distance traveled comprising: [Equation 9] performing [Equation 10] where x contains the reshape factor for the original window width (W) of the i-th frame, N contains the number of frames for calculating the reshape factor, dist contains the Euclidean distance, and x t The method of claim 10 , wherein x comprises the x coordinates of the skeleton points in frame t.
15. The determination of the average travel distance further comprises: [0011] performing [0012] where yi contains the reshape factor for the original window height (H) of the i-th frame, N contains the number of frames for calculating the reshape factor, dist contains the Euclidean distance, and yi t The method of claim 14 , wherein x comprises the y coordinates of skeleton points in the frame t.
16. said adjusting the size of said visual window comprises: [0013] where W′ comprises the reshaped window width and W comprises the original window width; [0014] The method of claim 10 , wherein {tilde over (x)} comprises a reshape factor for the original window width W of the i-th frame.
17. The adjusting of the size of the visual window further comprises: [Equation 15] where H' comprises the reshaped window height and H comprises the original window height; [0016] 17. The method of claim 16, wherein {tilde over (x)} comprises a reshape factor for the original window height H of the i-th frame.
18. upsampling, by said processor, said skeleton points to said multiple digital levels; The method of claim 10 further comprising:
19. providing at least one support service for at least one of generating, integrating, hosting, maintaining, and deploying computer-readable code in the hardware device, the code being executed by the processor to perform the receiving, the extracting of the skeleton points, the categorization, the generating of the initial visual window, the determining, the adjusting, the extracting of the feature vectors, the extracting of the point coordinates, the first linking, the generating of the convolutional neural network, and the enabling. The method of claim 10.
20. A computer program product causing a hardware device to carry out the method according to any one of claims 10 to 19.
Citation Information
Patent Citations
Neutral hand tracking, pose classification, and interface control
JP2013508827A
Program, device, and method for recognizing actions of persons using a plurality of recognition engines
JP2019144830A