Natural language video analytics system and method for processing one or more natural language video analytics commands
The natural language video analytics system addresses inefficiencies in conventional systems by using a neural network to process commands, improving user interaction and enabling real-time, error-reduced video analysis.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2026-03-13
AI Technical Summary
Conventional video analytics systems require technical expertise and are inefficient for users lacking training, leading to operational delays and increased error likelihood, especially in real-time environments.
A natural language video analytics system using a trained neural network to process commands, generate machine-readable instructions, and update displays, enhancing user interaction and real-time decision-making.
Enables efficient and effective real-time video analysis with reduced errors, allowing non-technical users to interact seamlessly and make informed decisions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Natural language video analytics system and method for processing one or more natural language video analytics commands technical field
[0001] The present invention relates globally to a natural language video analytics system and a method for processing one or more natural language video analytics commands. Prior art
[0002] Video analytics systems are widely used in various applications, such as retail analytics, traffic management, security, and surveillance. These systems use algorithms to process video data and extract insights and information. For example, in security applications, video analytics can detect suspicious behavior, identify individuals, and alert staff to potential threats. In retail environments, video analytics can help retailers understand customer behavior, optimize store layouts, and improve marketing strategies. Traffic management applications can benefit from real-time analysis of vehicle flow, traffic congestion detection, and incident management.
[0003] Traditionally, users operate video analytics systems via graphical user interfaces (GUIs) or command-line controls. These conventional methods can be complex and inefficient, especially for users lacking technical expertise. The processes often require extensive training and a thorough knowledge of the software, leading to operational delays and an increased likelihood of errors. Furthermore, the reliance on manual interaction restricts the scalability and responsiveness of these systems, particularly in environments requiring real-time analysis and rapid decision-making.
[0004] Consequently, there is a need for a natural language video analytics system and a method for processing one or more natural language video analytics commands that seek to resolve some of the aforementioned problems. Furthermore, other desirable functions and features will become apparent from the following detailed description and the attached claims, taken in conjunction with the accompanying drawings and the context of this disclosure. Summary of the invention
[0005] One aspect of this disclosure proposes a natural language video analytics system. The system includes at least one processor and at least one memory containing computer program code.At least one processor, at least one memory, and the computer program code are configured to enable the system to receive one or more natural language video analytics commands associated with video data, determine the validity of the one or more natural language commands using a trained neural network, in response to a positive determination of the validity of the one or more natural language commands, generate machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network, and transmit the machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.
[0006] To determine the validity of one or more natural language commands, the system can be configured to determine the relevance of one or more natural language video analytics commands for video data analytics using the trained neural network and, optionally, determine whether one or more natural language video analytics commands fall within the processing capability of the natural language video analytics system using the trained neural network.
[0007] Machine-readable video analytics instructions may include video overlay instructions, and the system may be configured to receive video data and video overlay instructions, modify video data based on video overlay instructions, and transmit the modified video data based on video overlay instructions to the output display.
[0008] Machine-readable video analytics instructions may include video transformation instructions, and the system may be configured to receive video data and video transformation instructions, modify video data based on video transformation instructions, and transmit the modified video data based on video transformation instructions to the output display.
[0009] The system can be configured to compare machine-readable video analytics instructions to an access control list, and in response to a comparison result indicating authorized access, retrieve video-derived data from one or more databases associated with the video dataset, based on the machine-readable video analytics instructions, and generate one or more data representations using the retrieved data and the machine-readable video analytics instructions, and transmit one or more data representations to the display module.
[0010] The system can be configured to generate a text response based on the retrieved data and machine-readable video analytics instructions, and transmit the text response to the display module. The system can also be configured to generate video-derived data based on a predetermined set of video analytics instructions using the video data and a trained video analytics algorithm, and store the video-derived data and the video data in one or more databases.
[0011] Another aspect of this disclosure proposes a method for processing one or more video analytics commands in natural language.The method includes receiving, by a processing device, one or more natural language video analytics commands associated with video data, determining, using the processing device, the validity of the one or more natural language commands with the help of a trained neural network, in response to a positive determination of the validity of the one or more natural language commands, generating, using the processing device, machine-readable video analytics instructions based on the one or more natural language commands with the help of the trained neural network, and transmitting, using the processing device, the machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.
[0012] The step of determining the validity of one or more commands in natural language using the trained neural network may include one or more of the steps of determining, using the processing device, the relevance of one or more video analytics commands in natural language for video data analytics using the trained neural network, and determining, using the processing device, whether the one or more video analytics commands in natural language fall within a processing capability of the video analytics system in natural language using the trained neural network.
[0013] Machine-readable video analytics instructions may include video overlay instructions, and the method may include the display module receiving video data and video overlay instructions, modifying the video data using the display module based on the video overlay instructions, and transmitting the modified video data based on the video overlay instructions to the output display using the display module.
[0014] Machine-readable video analytics instructions may include video transformation instructions, and the method may include the receipt, by the display module, of video data and video transformation instructions, the modification, using the display module, of video data based on video transformation instructions, and the transmission, using the display module, of the video data modified based on video transformation instructions to the output display.
[0015] The method may further include the steps of comparing, using the processing device, machine-readable video analytics instructions with an access control list, in response to a comparison result indicating authorized access; retrieving, using the processing device, video-derived data from one or more databases associated with the video data, based on the machine-readable video analytics instructions; generating, using the processing device, one or more data representations using the retrieved data and the machine-readable video analytics instructions; and transmitting, using the processing device, one or more data representations to the display module, the display module being configured to update the output display based on the one or more data representations.
[0016] The method may also include the steps of generating, using the processing device, a text response based on the retrieved data and machine-readable video analytics instructions, and transmitting, using the processing device, the text response to the display module. The method may further include the steps of generating, using the processing device, video-derived data based on a predetermined set of video analytics instructions using the video data and a trained video analytics algorithm, and storing the video-derived data and the video data in one or more databases. Brief description of the drawings
[0017] Embodiments of the invention will be better understood and readily apparent to a person skilled in the art from the following written description, given by way of example only, and in conjunction with the drawings, in which:
[0018] Fig. 1 shows a diagram of a natural language video analytics system, according to the disclosure embodiments.
[0019] Fig. 2 shows a diagram of an example implementation of the natural language video analytics system of Fig. 1, according to the disclosure embodiments.
[0020] Fig. 3 shows a flowchart illustrating a process for processing one or more natural language video analytics commands, according to the disclosure embodiments.
[0021] Fig. 4 shows a diagram of a computer device used to implement the system of Fig. 1.
[0022] Those skilled in the art will appreciate that the elements in the figures are illustrated in such a way as to maintain simplicity and clarity and that they have not necessarily been drawn to scale. For example, the dimensions of some of the elements in the illustrations, functional diagrams, or flowcharts may be exaggerated relative to other elements to improve understanding of the present embodiments. Description of the implementation methods
[0023] Embodiments of the present invention will be described, by way of example only, with reference to the drawings. Similar reference numbers and characters in the drawings refer to similar elements or equivalents. The detailed description that follows is purely illustrative and is not intended to limit the invention or its application and uses. Furthermore, there is no intention to be bound by any theory presented in the preceding context of the invention or in the detailed description that follows. A modular fluid treatment tank according to the embodiments shown herein is presented here, offering advantages of transportability, modularity, and scalability.
[0024] Certain parts of the following description are explicitly or implicitly presented in terms of algorithms and functional or symbolic representations of operations on data within computer memory. These algorithmic descriptions and functional or symbolic representations are the means used by those competent in the field of data processing to transmit the substance of their work as efficiently as possible to others who are also competent in the field. Here, and more generally, an algorithm is conceived as a coherent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities such as electrical, magnetic, or optical signals that can be stored, transferred, combined, compared, and otherwise manipulated.
[0025] Unless otherwise indicated, and as will be apparent from what follows, it will be understood that throughout this specification, discussions using terms such as "associate", "calculate", "compare", "determine", "retransmit", "generate", "identify", "include", "insert", "modify", "receive", "replace", "retrieve", "scan", "store", "transmit", or similar, refer to the action and processes of a computer system, or similar electronic device, which manipulates and transforms data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or any other device for storing, transmitting, or displaying information.
[0026] This specification also discloses an apparatus for carrying out the process operations. This apparatus may be specially constructed for the required purposes or may include a computer or other computing device that is selectively activated or reconfigured by a computer program stored within it. The algorithms and displays presented here are not intrinsically linked to any particular computer or other device. Various machines can be used with programs that comply with the teachings of the present invention. Alternatively, the construction of more specialized apparatus to carry out the required process steps may be appropriate. The structure of a computer will appear in the description below.
[0027] Furthermore, this specification also implicitly discloses a computer program, in the sense that it would be obvious to a person skilled in the art that the individual steps of the process described herein can be implemented by computer code. The computer program is not intended to be limited to any particular programming language and its implementation. It will be understood that a variety of programming languages and their coding methods may be used to implement the teachings of the disclosure contained herein. Moreover, the computer program is not intended to be limited to any particular command flow. There are many other variations of the computer program that may use different command flows without departing from the spirit or scope of the invention.
[0028] Furthermore, one or more steps of the computer program can be performed in parallel rather than sequentially. This computer program can be stored on any computer-readable medium. Computer-readable media can include storage devices such as magnetic or optical disks, memory chips, or other storage devices suitable for interfacing with a computer. Computer-readable media can also include wired media, as illustrated by the Internet system, or wireless media, as illustrated by the GSM mobile phone system. When loaded and executed on a computer, the computer program becomes a device that efficiently implements the steps of the preferred process.
[0029] In embodiments of the present invention, the term "server" may refer to a single computer device or at least a computer network of interconnected computing devices that work together to perform a particular function. In other words, the server can be contained within a single hardware unit or be distributed across several or many different hardware units.
[0030] The term "configured for" is used in the specification in relation to computer systems, devices, and program components. For a system consisting of one or more computers to be configured to perform particular operations or actions means that the system has been installed with software, firmware, computer hardware, or a combination thereof that, when in operation, causes the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operations or actions.When a specialized logic circuit assembly is configured to perform particular operations or actions, it means that the circuit assembly has electronic logic that performs the operations or actions.
[0031] The embodiments of this disclosure propose a natural language video analytics system and a method for processing one or more natural language video analytics commands. The natural language video analytics system may include a trained neural network, hereinafter referred to interchangeably as generative artificial intelligence (GenAI) or large language model (LLM).In the exemplary embodiments, the natural language video analytics system can receive one or more natural language video analytics commands associated with video data, generate machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network, and transmit the machine-readable video analytics instructions to a display module configured to update an output display based on the machine-readable video analytics instructions. In other words, the natural language video analytics system according to the embodiments of the invention can leverage trained neural networks to enhance video analysis and analytics capabilities, and can assist users in real-time decision-making, monitoring, and anomaly detection.
[0032] In the embodiments of this disclosure, the video analytics system can process and analyze video data to extract information. The system can use algorithms, including neural networks, to detect and interpret patterns, objects, and events within the data. Video. The system can be used in a wide range of applications, including but not limited to security monitoring, traffic monitoring, and behavioral analysis. The video analytics system, based on disclosure implementation methods, can provide more effective and efficient decision-making and real-time alerts based on analyzed video data.
[0033] In the exemplary embodiments, the natural language video analytics system can be configured to execute computer program code, hereinafter interchangeably referred to as one or more applications including, but not limited to, a video analytics application, a data visualization application and an image resolution enhancement application.In the exemplary embodiments, the video analytics application may include machine-readable instructions which, when executed by a processor of the natural language video analytics system, cause the processor to use the trained neural network to generate machine-readable instructions based on a user's natural language instructions, and to use the machine-readable instructions to dynamically adjust image processing parameters, enhance display settings, enable real-time analysis, and manage video sequences on the output display.
[0034] In the exemplary embodiments, the data visualization application may include machine-readable instructions which, when executed by the processor of the natural language video analytics system, cause the processor to use the trained neural network to generate machine-readable instructions based on natural language instructions from the user, and to generate data visualization information based on the machine-readable instructions and update the output display using the data visualization information. The data visualization application may also cause the natural language video analytics system to generate instructions that can be used to dynamically resize and / or rearrange the data visualizations presented on the output display to facilitate user analysis.
[0035] The embodiments serving as examples in this disclosure may also include the image resolution enhancement application comprising computer-readable instructions which, when executed by the processor of the natural language video analytics system, cause the processor to increase the resolution of an image beyond its original resolution using a general adversarial network (GAN). A suitable general adversarial network (GAN) may include, but is not limited to, Real-World Enhanced Super-Resolution General Adversarial Network (ESRGAN). (Enhanced in the Real World). The image resolution enhancement application can also enable the natural language video analytics system to estimate the number of individuals in a designated area of an image (i.e., to count the number of people forming a crowd) and perform crowd counting using HRNet (High-Resolution Net) to generate segmentation maps and FIDT (Focal Inverse Distance Transform) to locate crowds on the generated segmentation maps. In some embodiments, an image can be a frame within a sequence of frames which, when played back sequentially, form a video.
[0036] In the exemplary embodiments, video data may include, but is not limited to, one or more of the following: live video stream data and recorded video data, where live video stream data is video content captured and transmitted in real time, and recorded video data is video content captured and stored for later playback or analysis. Video data may include a digital representation of content captured by one or more image capture devices and may be stored as image sequences or frames. Video data may also include metadata such as resolution, frame rate, encoding format, duration, and additional information such as timestamps, camera settings, and geolocation data.In embodiments of the invention, video data may include video analytics data, video analytics data being information derived from the analysis of video data using the processing described below, and may include, but not be limited to, patterns, trends, and metrics derived from video data.
[0037] In the exemplary embodiments, the video analytics instructions may include, but are not limited to, one or more of the following: video overlay instructions and video transformation instructions. The video overlay instructions may include machine-readable instructions that can cause the video analytics system to include additional visual elements such as text, images, or graphics on a video displayed on an output display. The instructions may also include instructions related to the position of the overlaid elements relative to the video displayed on the output display, the interaction of the overlaid elements with the underlying video content, and the conditions for displaying or modifying the overlays.The elements may include, but are not limited to, data representations associated with video analytics, character matrices, tracking identifiers, frames per. frame rate (FPS, frames per second) and regions of interest (ROI, regions of interest) are used to measure video transformation. Video transformation instructions can include machine-readable instructions that cause the video analytics system to modify or manipulate the video displayed on an output display. These instructions may include, but are not limited to, resizing, cropping, rotating, color-balancing, filtering, or applying effects to the video displayed on an output display.
[0038] Figure 1 shows a diagram of a natural language video analytics system 100, according to the disclosure embodiments. In the exemplary embodiments, the system 100 may include at least one processor 102 and at least one memory 104 containing computer program code. The at least one processor 102 and the at least one memory 104 may be housed in a server 106. The server 106 may also include a display module 108 configured to refresh an output display.At least one processor 102, at least one memory 104, and the computer program code are configured to enable the system 100 to receive one or more natural language video analytics commands associated with video data, determine the validity of the one or more natural language commands using a trained neural network, in response to a positive determination of the validity of the one or more natural language commands, generate machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network, and transmit the machine-readable video analytics instructions to the display module 108, the display module 108 being configured to update the output display based on the machine-readable video analytics instructions.
[0039] In embodiments of this disclosure, the display module 108 may be a separate server, or a component or subsystem within the server 106 that can render and present visual information to a user. The display module may include, but is not limited to, a graphics processing unit configured to render images, videos, and other graphical data, an output display, and a circuit assembly configured to present visual information to the user. The display module is configured to facilitate interaction between the user and the server associated with the display module by converting electronic signals into visual information.
[0040] Figure 2 shows a diagram of an example implementation of the Natural Language Video Analytics System 100 of Figure 1, according to the disclosure embodiments. The Natural Language Video Analytics System 100 can be configured to run a Video Analytics Application 202, hereinafter interchangeably referred to as the Video Analytics (VA) Copilot Module, associated with various video analytics functions. As will be described in detail below, the video analytics application 202 according to embodiments of the invention may include a set of instructions in a machine-readable format that can be executed by the natural language video analytics system 100 to perform the various video analytics functions described herein. In the exemplary embodiments, the video analytics application 202 can cause the natural language video analytics system 100 to generate machine-readable video analytics instructions based on one or more natural language commands from a user using the trained neural network.That is to say, the System 100 can generate video analytics instructions based on natural language user commands, and these video analytics instructions can be associated with delayed video playback modifications that can dynamically adjust image processing parameters, improve display settings, enable real-time analysis, and manage video sequences.
[0041] The video analytics application 202, according to the disclosure embodiments, may include one or more handlers, each being a subroutine or program configured to perform one or more specific functions within the application. Each of the one or more handlers may include instructions in a machine-readable format that can be executed by the natural language video analytics system 100 to perform the various functions described in more detail below. In one example embodiment, the video analytics application 202 may include, but is not limited to, a chatbot handler 202a, a prompt classification handler 202b, a drawing handler 202c, a recording handler 202d, and a multi-video viewing handler 202e.In the embodiments of this disclosure, the video analytics application 202 may cause the natural language video analytics system 100 to receive one or more natural language video analytics commands associated with video data, determine the validity of the one or more natural language commands using a trained neural network, generate machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network in response to a positive determination of the validity of the one or more natural language commands, and transmit the machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.By determining the validity of one or more natural language commands, the video analytics application 202 can lead the natural language video analytics system 100 to determine the relevance of one or more natural language video analytics commands. natural for video data analytics using the trained neural network and optionally, to determine whether one or more natural language video analytics commands fall within a processing capability of the natural language video analytics system using the trained neural network.
[0042] In embodiments of this disclosure, machine-readable video analytics instructions may include video overlay instructions, and the video analytics application 202 may cause the natural language video analytics system 100 to receive the video data and video overlay instructions, modify the video data based on the video overlay instructions, and transmit the modified video data based on the video overlay instructions to the output display.
[0043] In exemplary embodiments, machine-readable video analytics instructions may include video transformation instructions, and the video analytics application 202 may cause the natural language video analytics system 100 to receive the video data and the video transformation instructions, modify the video data on the basis of the video transformation instructions and transmit the modified video data on the basis of the video transformation instructions to the output display.
[0044] The video analytics application 202 can also cause the natural language video analytics system 100 to compare machine-readable video analytics instructions to an access control list, retrieve video-derived data from one or more databases associated with the video dataset, based on machine-readable video analytics instructions in response to a comparison result indicating authorized access, generate one or more data representations using the retrieved data and machine-readable video analytics instructions, and transmit the one or more data representations to the display module.The 202 video analytics application can also cause the 100 natural language video analytics system to generate video-derived data based on a predetermined set of video analytics instructions using video data and a trained video analytics algorithm, and store the video-derived data and video data in one or more databases.
[0045] In the embodiments of this disclosure, the chatbot manager 202a can cause the video analytics system 100 to generate instructions associated with managing a graphical user interface (GUI) and to facilitate user interaction between the user and the natural language video analytics system 100. For example, the chatbot manager 202a can cause the video analytics system 100 to receive one or more User commands, execute a prompt classification manager 202b for further processing of one or more user commands, and display a response indicating the processing result via the GUI. The chatbot manager 202a can also cause the video analytics system 100 to generate feedback messages using the drawing manager 202c and the recording manager 202d, and display the feedback messages via the GUI. In the embodiments of this disclosure, the drawing manager 202c can cause the system 100 to perform drawing operations on video frames, while the recording manager 202d can cause the system 100 to retrieve recordings and metadata.
[0046] The prompt classification manager 202b can cause the video analytics system 100 to transmit and receive messages from a large language model (LLM), i.e., to maintain an LLM session. The prompt classification manager 202b can cause the video analytics system 100 to launch the LLM session with a pre-written prompt containing instructions for processing user requests, and to use the LLM to perform one or more of the following: (i) relevance assessment, (ii) feasibility analysis, and (iii) requirements categorization.In one embodiment, the prompt classification manager 202b, in a relevance evaluation, can cause the video analytics system 100 to determine the relevance of one or more natural language video analytics commands for analyzing the video dataset using the trained neural network based on a user-defined filtering prompt. In one embodiment, the user-defined filtering prompt can include contextual information about a video analytics tool, such as keywords, examples, and guidelines sufficiently relevant for the LLM to determine whether the commands are related to a video analytics tool. In one embodiment, any command deemed not to fall within the scope of a video analytics tool is classified as "irrelevant."The prompt classification manager 202b can cause the video analytics system 100 to send a message indicating the determination result and to execute the chatbot manager 202a, which can cause the system 100 to display a text explanation to the user via the GUI. If the user request is deemed relevant for video analytics, it is classified as "relevant".
[0047] In a feasibility analysis, the prompt classification manager 202b can enable the video analytics system 100 to determine, using the trained neural network, whether one or more natural language video analytics commands fall within the processing capabilities of the natural language video analytics system 100. In one example, the feasibility of a command can be determined based on the capabilities of the natural language video analytics system and the type of data it can process, for example, in association with character matrices, tracking identifiers, and a region of interest (ROI). If the command is deemed impractical, the prompt classification manager 202b can cause the video analytics system 100 to execute the chatbot manager 202 to display a message indicating the result of the feasibility determination to the user via the GUI. In one example embodiment, the feasibility analysis of the natural language video analytics command can follow a relevance assessment of the natural language video analytics command.
[0048] In the requirements categorization process, the prompt classification manager 202b can cause the video analytics system 100 to categorize commands into one of two categories: display or recording. Commands that require real-time video stream processing are classified in the "display" category and will be processed by the drawing manager 202c. Commands that involve processed or stored video data are classified in the "recording" category and will be processed by the recording manager 202d.
[0049] In embodiments of this disclosure, video data may include, but are not limited to, video stream data, and machine-readable video analytics instructions generated based on one or more natural language commands using the trained neural network may include video overlay instructions. In the exemplary embodiments, the video overlay instructions may include machine-readable instructions for drawing on real-time video data using features including, but not limited to, character matrices, tracking identifiers, frames per second (FPS), and regions of interest (ROI).The Natural Language Video Analytics System 100 can be configured to receive video data and video overlay instructions, modify the video data based on the video overlay instructions, and transmit the modified video data to the output display. In one embodiment, the Drawing Manager 202c can cause the Video Analytics System 100 to process the aforementioned steps and generate machine-readable video analytics instructions based on one or more natural language commands using the trained neural network process. In example embodiments, the Drawing Manager 202c can cause the Video Analytics System 100 to perform drawing operations on video data in real time using features such as character matrices, tracking IDs, frames per second (FPS), and regions of interest (ROI).As will be explained in the paragraph below, the manager. The 202c drawing manager can initialize the 100 video analytics system with predefined configurations and maintain a live display using the display module (e.g., showing live video feeds and associated real-time video analytics information). The 202c drawing manager can also enable the 100 video analytics system to use segmentation to process one or more natural language commands using the trained neural network without interrupting the video processing loop.
[0050] In exemplary embodiments, the drawing manager 202c can cause the natural language video analytics system 100 to execute a sequence of steps to overlay features such as character matrices, tracking identifiers, frames per second (FPS), and regions of interest (ROI) onto real-time video data shown on the output display. The steps include, but are not limited to, (i) receiving a message including one or more natural language commands, refresh processes, background information, and task constraints, (ii) generating machine-readable video overlay instructions (e.g.(a Python code) based on one or more natural language commands using the trained neural network, (iii) the execution of machine-readable video overlay instructions, (iii) the modification of video data based on the video overlay instructions, and (iv) the transmission of the modified video data to the output display. The drawing manager 202c can cause the video analytics system 100 to generate a feedback message associated with the above processing, and the feedback message can be accessed by the chatbot manager 202a via a shared memory segment on the natural language video analytics system 100.
[0051] In embodiments of this disclosure, "update methods" include specific actions or techniques that dictate how requested changes are to be applied to the video. The methods include, but are not limited to, user requests to modify display settings such as text and background colors, toggle the visibility of elements, and adjust filtering preferences. "Background information" includes contextual information provided by the user for the drawing manager to produce the desired output. For example, background information might include details about the video stream format, existing overlays, the current lag playback speed, or specific areas of interest within the video.Contextual information can ensure that the generated instructions are appropriate and relevant to the current state of the video data. "Task constraints" include limitations or requirements that the drawing manager must respect while processing the user request. Limitations or requirements may include, but are not limited to, specific regions where overlays should not be applied or compatibility requirements with existing video data formats. These limitations or requirements can ensure that the generated code is feasible and aligned with the system's capabilities and needs.
[0052] In embodiments of this disclosure, the video data may include, but is not limited to, one or more sets of recorded video stream data, and the machine-readable video analytics instructions generated based on one or more natural language commands using the trained neural network may include video transformation instructions. The natural language video analytics system 100 can be configured to receive the video data and the video transformation instructions, modify the video data based on the video transformation instructions, and transmit the modified video data to the output display. In one embodiment, the recording manager 202d can cause the video analytics system 100 to perform the aforementioned steps.The 202d recording manager can enable the 100 video analytics system to generate machine-readable video transformation instructions based on one or more natural language commands using the trained neural network process.
[0053] In embodiments, the record manager 202d can cause the natural language video analytics system 100 to generate video-derived data based on a predetermined set of video analytics instructions using the video data and a trained video analytics algorithm, and store the video-derived data and the video data in one or more databases. The video-derived data can include video analytics data associated with the video data.In one example embodiment, the 202d recording manager can cause the 100 natural language video analytics system to retrieve video data and video-derived data stored in one or more databases based on machine-readable video analytics instructions; generate one or more data representations using the retrieved data and the machine-readable video analytics instructions; and transmit the one or more data representations to the display module. The one or more data representations may include, but are not limited to, additional video analytics data and statistical information associated with the video data, video-derived data, or both video data and video-derived data.
[0054] In one embodiment, the recording manager 202d can cause the natural language video analytics system 100 to (i) receive a message including one or more natural language commands, information In the background, the refresh processes, arguments (i.e., any additional input requested by the LLM), and task constraints, (ii) launch a new process to execute the code or return response, and (iii) share the feedback message with the Chatbot Manager via shared memory. In some embodiments, the return response may include a "processing_record" process for later execution.
[0055] In the embodiment examples, the multi-video viewer manager 202e can cause the video analytics system 100 to display multiple videos side-by-side simultaneously on an output display. The multiple videos can include, but are not limited to, an original video and a video modified by the natural language video analytics system 100 based on one or more natural language video analytics commands. The multi-video viewer manager 202e can also cause the video analytics system 100 to receive and process natural language video analytics commands associated with video window management, the video windows being graphical user interface elements that display video content on the output display. Each window can show a separate video stream or file and can be individually controlled, resized, and repositioned by the user.
[0056] In the example embodiments, a security manager (not shown) can cause the natural language video analytics system 100 to execute instructions that can enhance the security of the natural language video analytics system 100. In one example embodiment, access to the video analytics application 202 may require users to first perform multi-factor authentication. Furthermore, access levels may be based on user roles. For example, an operator of the natural language video analytics system 100 may access basic functionalities such as live stream management, reviewing stored recordings, and performing system maintenance tasks, while having limited control over user management and system settings.A viewer or guest may have restricted access to read live streams only and cannot review stored recordings or make changes to the system. An administrator has full access to system features, settings, and data, and can manage users, permissions, and configurations of the Natural Language Video Analytics 100 system.
[0057] The video analytics application 202 according to the embodiments of this disclosure can resolve the technical problems associated with the secure execution of video analytics queries, the generation of efficient machine-readable video analytics instructions, and the generation of a user interface for navigation and analysis. Advantageously, the video analytics application 202 according to the embodiments of this disclosure can (i) generate machine-readable video analytics instructions from data based on users' natural language commands, (ii) generate and implement a video analytics data visualization, i.e., generate a suitable data visualization and implement the necessary modifications to the video data to properly represent the video analytics data, (iii) generate an effective user interface, i.e., generate a display of visual information where users can interact with the displayed data and manipulate the dashboard components as needed (e.g.(iv) a user interface that can facilitate navigation of different components, i.e. a chat assistant, a video window, record, play, pause and save buttons, a history file, etc., and that facilitates analysis, i.e., presenting a multi-video window for comparison, the ability to rearrange software gadgets, an analytical summary console, alerts, etc.) and (iv) strengthen security measures to prevent malicious attacks, i.e. implement security measures to reduce the number of unauthorized accesses to video data.
[0058] Embodiments of this disclosure also propose a data visualization application 204, hereinafter interchangeably referred to as the dashboard co-pilot module. The data visualization application 204 can cause the video analytics system 100 to use the trained neural network to generate machine-readable instructions based on natural language instructions from the user, generate data visualization information based on the machine-readable instructions, and update the output display using the data visualization information. The data visualization application 204 can cause the video analytics system 100 to generate instructions that can dynamically resize and / or rearrange the data visualizations presented on the output display to facilitate user analysis.
[0059] In the exemplary embodiments, the data visualization application 204 may include a text input module 204a that can cause the video analytics system 100 to receive one or more natural language video analytics commands associated with video data. The one or more natural language video analytics commands may include, but are not limited to, information to be extracted from the video data and optionally, the method or format used to graphically represent the data (e.g., bar charts, line graphs, etc.). The data visualization application 204 may also include a database management command generation component (e.g. a Structured Query Language (SQL) command) 204b which can cause the video analytics system 100 to determine the validity of one or more natural language commands using a trained neural network, and generate machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network in response to a positive determination of the validity of the one or more natural language commands.In one embodiment, the data visualization application 204 can cause the video analytics system 100 to generate a prompt for a Language Lead Management (LLM) based on one or more natural language video analytics commands, and to generate an SQL query using the LLM based on the prompt. The prompt includes information about the database associated with the video-derived data and the video data, example queries and responses, and additional instructions that may include renaming column headers from a results table associated with a response to the SQL query. The generated SQL query can be used to retrieve information from the database, where the results are stored as a table containing the renamed column headers and values.
[0060] In the exemplary embodiments, the data visualization application 204 may include a security component 204c. The security component 204c may cause the video analytics system 100 to execute a sequence of steps to determine the relevance of the output of the text input module 204a, compare the machine-readable video analytics instructions to an access control list, and authorize further processing of the machine-readable video analytics instructions in response to a comparison result indicating authorized access. In one exemplary embodiment, the security component 204c may cause the video analytics system 100 to receive a credential for verification and authentication from a user before allowing the user access to the data visualization application 204.The authenticator can also determine the database privileges and permissions available to the user. In one embodiment, the generated SQL query can be subjected to one or more of a blocklist check and an allowlist check to reduce the number of SQL injection attacks. The blocklist can be used to reject SQL queries that include forbidden commands or keywords. The allowlist can be used to reject SQL queries that do not include certain commands or keywords. In one embodiment, only queries that do not have... The entries rejected by these two lists will be used to extract information from the database.
[0061] In the exemplary embodiments, the data visualization application 204 may include an analytics generation module 204d. The analytics generation module 204d may cause the video analytics system 100 to execute a sequence of steps to generate a text response to the command in natural language based on data extracted using an SQL query generated by component 204b with the prompt. The prompt may include the data extracted using the SQL query generated by component 204b and additional instructions.
[0062] The data visualization application 204 may also include a plot generation module 204e. The plot generation module 204e may cause the video analytics system 100 to generate one or more data representations, hereinafter also referred to as data visualization, using the results table with renamed column headers and a prompt using the LLM. The prompt may include the results table with renamed column headers, examples of plot types to use depending on the query, and additional instructions. The output is Python code that will be executed to perform the interactive data visualization using the Plotly library.
[0063] The data visualization application 204 may also include a dashboard module 204f. The dashboard module 204f can enable the video analytics system 100 to create a dashboard on a user interface (UI) using the Dash and Dash Draggable libraries. The four buttons are created using the Dash library, while the rest of the dashboard is created using the Dash Draggable library.
[0064] The data visualization application 204 according to the embodiments of this disclosure can solve technical problems associated with the secure execution of database queries, efficient data visualization and the creation of a dynamic and interactive dashboard.The data visualization application 204 according to the embodiments of this disclosure can (i) generate accurate machine-readable data visualization instructions based on users' natural language commands, i.e., understand user requirements, determine appropriate data to extract and the optimal plot type for data representation, (ii) generate database SQL queries, i.e., generate machine-readable database query instructions to facilitate processing and retrieval of relevant data from a database, (iii) generate a data visualization, i.e., generate a suitable data visualization. to properly represent the data using a corresponding plotting library, e.g. Python and Plotly respectively, (iv) generate a dynamic and interactive dashboard, i.e. generate a display of visual information where users can interact with the displayed data and manipulate the dashboard components as needed, and (v) strengthen security measures to prevent malicious attacks, i.e. implement security measures to reduce the number of malicious attacks such as an SQL injection attack on the database and a DDoS attack on the application.
[0065] Embodiments of this disclosure also propose an image resolution enhancement application 206, hereinafter referred to interchangeably as the super-resolution module. The image resolution enhancement application 206 can cause the video analytics system 100 to increase the resolution of an image beyond its original resolution using a general adversarial network (GAN). The image resolution enhancement application 206 according to the disclosure embodiments may include one or more handlers, each being a subroutine or program configured to perform one or more specific functions within the application. Each of the one or more handlers may include instructions in a machine-readable format that can be executed by the natural language video analytics system 100 to perform the various functions described in more detail below.In one example embodiment, the image resolution enhancement application 206 may include, but is not limited to, an image loader manager 206a, a scaling manager 206b, a super-resolution manager 206c, a drive manager 206d, a drawing manager 206e, a crowd counting manager 206f, and a multi-image viewing manager 206g.
[0066] In embodiments of this disclosure, the image loader manager 206a can cause the video analytics system 100 to process one or more images and transmit the one or more images to the output display. In one embodiment, when a plurality of images are processed, the video analytics system 100 can display all the images simultaneously and can be configured to receive one or more user commands to select one or more images to upscale, add or delete images, edit filenames, rearrange files, view file metadata, and preview images.
[0067] In the exemplary embodiments, the scaling manager 206b can cause the video analytics system 100 to receive input from the user, the input being associated with a scaling factor for a display to the next higher resolution of the user-selected images. The user may be offered default options to perform upscaling by a factor of 2 or 4. Alternatively, users can also provide their own pre-trained models with their own defined degree of upscaling. In the example implementations, the 206c super-resolution manager can cause the 100 video analytics system to increase the resolution of a user-selected image beyond its original resolution using a general adversarial network (GAN). In one example implementation, a suitable general adversarial network (GAN) can be used. The GAN can include, but is not limited to, Real-ESRGAN (Real-World Enhanced Super-Resolution General Adversarial Network) by the OpenMMLab 2.0 library. The model is based on the PyTorch application framework.
[0068] In the exemplary embodiments, the Image Resolution Enhancement Application 206 may include the Training Manager 206d. The Training Manager 206d can enable the Video Analytics System 100 to facilitate the training and storage of user-customized super-resolution models for specific datasets. In one embodiment, the user-generated training dataset may be divided into two folders: a source folder, containing the images to be upscaled, and a target folder, containing the corresponding upscaled images in the same order and with the same filenames as the source folder. The scaling factor for the upscaling process may be specified by the user.When the training process is initiated by the 100 Video Analytics system, a progress bar can be displayed to indicate the training progress. A line of text can be shown below the progress bar to display the quality metrics used to evaluate training progress. These metrics include, but are not limited to, the peak signal-to-noise ratio (PSNR), which measures the quality of the super-resolution images relative to the original images. The 206d Training Manager can save and overwrite the model that achieves the best PSNR. The user may have the option to terminate the training process if the quality is deemed unsatisfactory. In such cases, either the model can be saved in its current state, or the automatically saved model can be used.Once the training is completed or has been terminated, a text file documenting the training progress can be saved in the same directory as the model.
[0069] In the exemplary embodiments, the Image Resolution Enhancement Application 206 may include the Drawing Manager 206e. The manager Drawing Manager 206e allows the Video Analytics 100 system to provide drawing tools for image annotation within the multi-image viewer. Customizations, including the thickness and color of the drawing tool, are also available through Drawing Manager 206e. An undo function facilitates returning to previous annotations. Additionally, an option is provided to save images with or without applied annotations.
[0070] In the exemplary embodiments, the image resolution enhancement application 206 may include the crowd counting manager 206f. The crowd counting manager 206f may cause the video analytics system 100 to execute a sequence of steps to count the number of people forming a crowd in images. The sequence includes (i) the use of a pre-trained High Resolution Network (HRNet) to perform semantic segmentation on the image, so that individuals are segmented from the foreground, and (ii) the use of the Focal Inverse Distance Transform (FIDT) to obtain crowd localization within the segmentation map, where each individual is represented as a spot.The process also includes (iii) converting the segmentation map into a density map, which is then rendered so that it can be viewed within the multi-image viewer, and (iv) enumerating the number of maximum intensity points within each localized spot to count the number of people forming a crowd. The resulting number can be displayed as text overlaid on the density map.
[0071] In the exemplary embodiments, the Image Resolution Enhancement Application 206 may include the Multi-Image Viewer Manager 206g. The Multi-Image Viewer Manager 206g can enable the Video Analytics System 100 to facilitate the creation and viewing of multiple image windows that can be freely resized and rearranged on the output display. Each image window may include a drop-down list for selecting the image to be displayed. In addition, each image window may be equipped with its own drawing manager and its own crowd counting manager. The crowd counting manager may be excluded for the generated density maps.
[0072] In embodiments of this disclosure, the image resolution enhancement application 206 may advantageously provide an interface that facilitates close-up and side-by-side comparisons between the original and upscaled images, and provides tools for comprehensive image analysis and download history verification. In the example embodiments, the image resolution enhancement application 206 can increase the resolution of an image beyond its original resolution using a general adversarial network. (GAN) with a graphics processing unit (GPU, graphical processing unit) to perform fast inference.
[0073] Figure 3 shows a flowchart illustrating a method 300 for processing one or more natural language video analytics commands, according to the embodiments of the disclosure. The method 300 can be implemented by the server 106 of the system 100, hereinafter referred to interchangeably as the processing device.The method 300 includes overall step 302 of receiving, by a processing device, one or more natural language video analytics commands associated with video data, step 304 of determining, using the processing device, the validity of one or more natural language commands using a trained neural network, in response to a positive determination of the validity of one or more natural language commands, step 306 of generating, using the processing device, machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network, and step 308 of transmitting, using the processing device, the machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.
[0074] Embodiments of this disclosure also provide a graphical user interface (GUI) application. The GUI application can integrate the three aforementioned video analytics applications 202, data visualization application 204, and image resolution enhancement application 206 to facilitate efficient use and navigation. The following paragraphs describe each of the video analytics application 202, data visualization application 204, and image resolution enhancement application 206 in more detail and explain how the applications can interact with the GUI application.In some embodiments, the video analytics application 202 can enable the video analytics system 100 to provide a user-friendly chat assistant that simplifies interactions with a native language interface. Users can input their queries in natural language via the keyboard, which is then processed by the system 100 to reflect the requested changes directly onto the video stream displayed on the output screen. In some embodiments, the chat interface can also maintain the complete conversation history and provide responses in both text and graphical formats.In the back-end program, the 202 video analytics application can serve as an assistant that ensures user requests are handled securely and efficiently, using the trained neural network to process requests and generate the relevant code that will be seamlessly integrated into the GUI application. For example, . The Video Analytics 202 application can enable the Video Analytics 100 system to filter each user request based on security and usability, rejecting those that do not meet security standards with text feedback sent to the user via the chat interface. The Video Analytics 202 application can enable the Video Analytics 100 system to process user requests to modify display configurations, including, but not limited to, changes to background settings, toggling element visibility, and adjusting filtering preferences. The Video Analytics 202 application can enable the Video Analytics 100 system to perform real-time analysis, such as tracking and analyzing people's movements or filtering specific attributes within a crowd.By accessing the history of video recordings, the 202 video analytics application can enable the 100 video analytics system to retrieve specific video segments that are triggered by events (e.g., people entering or leaving, objects left behind, etc.) and provides users with the flexibility to play, edit, and save these extracted recordings.
[0075] In the exemplary embodiments, the Video Analytics Application 202, together with the GUI application, can enable the Video Analytics System 100 to facilitate user navigation, support multiple functionalities such as multi-video viewing and a prompt history file, and allow users to customize the on-screen interface, such as rearranging software gadgets or panels to display relevant information like live video feeds, a prompt-generated video compilation, analytical summaries, alerts, recent activity, etc. Therefore, the Video Analytics Application 202, according to the embodiments, can allow users without technical expertise to easily customize their video analytics and decision-making processes.
[0076] In the exemplary embodiments, the data visualization application 204 can be a web application. The data visualization application 204 can enable the video analytics system 100 to perform data visualization based on text commands from users, written in natural language, using a trained neural network. The neural network can be trained to generate an SQL query based on the user's text input and the corresponding connected database. The query can then be passed to the database to extract the relevant information in the form of a data frame. The data in the table can be presented as a table or a graph, depending on what is most suitable and the user's commands. For example, users can request Specific statistics and chart types are used in their query, which will be reflected in the generated graphical component. The Data Visualization 204 application can cause the Video Analytics 100 system to present information as individual graphical components on the output display, and these components can be dynamically resized and rearranged to facilitate analysis. To enable seamless real-time data analysis, the Data Visualization 204 application can cause the Video Analytics 100 system to continuously refresh the displayed data in real time to provide immediate insights and support decision-making. Furthermore, the Data Visualization 204 application can enable the Video Analytics 100 system to facilitate drag-and-drop of dynamically generated graphical components into the permanent static UI.
[0077] In the illustrative embodiments, the Image Resolution Enhancement Application 206 can cause the Video Analytics System 100 to enhance an image beyond its original resolution. The Image Resolution Enhancement Application 206 can cause the Video Analytics System 100 to accept multiple image formats as input and to receive a user-specified degree of upscaling or enlargement (e.g., multiplied by 2, 4, and 8) for the image upscaling process. The Image Resolution Enhancement Application 206 can cause the Video Analytics System 100 to generate higher-resolution images and can store the images in various image formats. The Image Resolution Enhancement Application 206 features a crowd counting function.The UI makes it easy for the user to upload input images, view the output in super-resolution, generate a corresponding density map, and display the resulting number of people. The Image Resolution Enhancement 206 application, together with the UI application, can enable the Video Analytics 100 system to overlay a density map for easier image analysis and maintain a scroll bar for upload history, which can be displayed on the output screen with a double-click. The Image Resolution Enhancement 206 application allows users to conveniently perform super-resolution, crowd counting, and image analysis within the same interface.
[0078] In the exemplary embodiments, the GUI application can be implemented as a web interface that integrates the aforementioned applications and facilitates secure access to these functionalities. The web interface can ensure that the applications are accessed in a unified and secure manner and can provide a platform for user interaction with the described applications. In the embodiments, each user can be associated with a distinct set of authentication methods, including different levels of access control and permissions for the features of each application. For example, a user assigned minimal viewing permissions will be limited to viewing the live video stream within the video analytics application 202 and the pre-generated real-time data visualizations within the data visualization application 204. In one example embodiment, the user will not gain access to the image resolution enhancement application 206, nor will they have the ability to perform any additional actions within the aforementioned video analytics and data visualization applications 202 and 204.
[0079] In one embodiment, the video analytics application 202 and the GUI application can enable the natural language video analytics system 100 to receive and display video analytics within an interactive dashboard. The functionality can be provided via a web application. The video analytics application 202 and the GUI application can enable the natural language video analytics system 100 to receive user input (e.g., one or more natural language video analytics commands) via a chat interface.The request can be processed by a prompt classification manager 202b as described above, which can cause the natural language video analytics system 100 to evaluate the user input by first determining whether the request is relevant to the software's capabilities, then the feasibility of the request in relation to the software's current functional state, and finally categorizing the request as either "Display" or "Record." The user can be notified if the input is deemed irrelevant or unfeasible. An input classified as "Display" can be processed by the drawing manager 202c, which can cause the video analytics system 100 to perform drawing operations on the video frames using real-time data features such as character matrices, a tracking identity (tracking ID), frames per second (FPS), and a region of interest (ROI).An example of user input might include, but is not limited to, "change the character matrix color to red." Input that is classified as a "record" is processed by the 202d Record Manager, which can then prompt the 100 Video Analytics System to retrieve and process the relevant records and metadata. An example of user input might include, but is not limited to, "count the number of people entering the store from 5:30 PM to 6:30 PM on August 15, 2023." The 202 Video Analytics application can then prompt the 100 Natural Language Video Analytics System to generate machine-readable video analytics instructions based on the natural language commands using a neural network. trained to display information associated with machine-readable video analytics instructions on an output display. These changes can then be viewed in the multi-video window manager on the user interface, where side-by-side comparisons between the original video window and the generated video windows can be performed. Automated analytical summaries, such as text overlaid on the generated video windows or a separate window on the interface, can also be generated. The generated videos can be stored in a widely supported format (e.g., H.264 / H.265 MP4, MOV, etc.).
[0080] In one embodiment, the data visualization application 204 and the GUI application can cause the natural language video analytics system 100 to generate a data visualization within an interactive dashboard. The functionality can be provided via a web application, which would require user authentication and verification before access is granted. The type of user privileges and permissions granted to the account will depend on the role assigned to the user. The data visualization application 204 and the GUI application can cause the natural language video analytics system 100 to receive a text query (e.g., video analytics instructions) in natural language from a user. The data visualization application 204 can cause the natural language video analytics system 100 to generate an SQL query using an LLM.In one embodiment, the text query can be passed to the LLM as a prompt containing a schema and a description of the database table to generate the SQL query with renamed column headers for user understanding. The generated query can be passed through one or more of the following security checks: a reject list and an allow list (i.e., the filter can be a reject list, an allow list, or both) to reduce the risk of SQL injection attacks. The filtered query can then be used to retrieve relevant results as a table with values and column headers from a PostgreSQL database.The extracted data table with renamed column headers can be passed to the LLM as a prompt to generate a text response to the text query. The same data table can also be passed to the LLM as a prompt to generate Python code for data visualization using the Plotly Python library. The generated code can be executed, and the data visualization output can be displayed on a user interface dashboard created using the Dash and Dash Draggable Python libraries. The output can be a single, movable graphical component that can be dragged and resized within the dashboard itself. It is also possible to generate multiple graphical components that can be freely rearranged on the dashboard for comparison. Users can save the graphical components to the dashboard(s), which can then be edited for one or more dashboards. In one example embodiment, two buttons can be implemented on the dashboard to allow users to perform the following actions: (i) a "Run Query" button to view the query as a removable graphical component (alternatively, pressing the "Enter" or "Return" key performs the same action, and pressing the button again will generate additional removable graphical components on the dashboard), and (ii) a "Clear" button to remove all graphical components from the dashboard.All generated graphical components can be freely rearranged and resized within the dashboard itself, making it a dynamic dashboard. Using graphics libraries such as the Plotly library for data visualization allows for interactive plots where users can hover over the plot to view additional information.
[0081] In one embodiment, the Image Resolution Enhancement Application 206 and the GUI application can cause the Natural Language Video Analytics System 100 to increase the resolution of an image beyond its original resolution using a General Adversarial Network (GAN). The functionality can be provided via a web application. The Image Resolution Enhancement Application 206 and the GUI application can cause the Natural Language Video Analytics System 100 to receive at least one image for resolution enhancement. In one embodiment, an option to upscale the image resolution by a factor of 2 or 4 can be presented to the user. The user may have an option to specify the desired output image format for saving the upscaled images.Once the upscaling process is complete, the Natural Language Video Analytics 100 system can display a multi-image window on the output display to facilitate side-by-side comparison between the original image and the upscaled images, with functionality provided to zoom in and out of the images. In one example embodiment, a drawing tool can be provided to allow annotations on the images. An option to save these changes is also available. If a "crowd count" request is received, the Natural Language Video Analytics 100 system can generate an additional density map image along with the corresponding number of people, presented as text overlaid within the multi-image window. The user can choose to save either the density map image, the upscaled images, or both. The user interface can include a dropdown menu that displays the entire download history, making it easier to identify duplicate images. In the example embodiments, the Natural Language Video Analytics System 100 can train and save custom super-resolution models using custom image datasets. The UI design can facilitate super-resolution upscaling and side-by-side comparisons.
[0082] In one embodiment, the GUI application can integrate the three aforementioned video analytics application 202, data visualization application 204 and image resolution enhancement application 206 to facilitate efficient use and navigation, and enable the efficient performance of various operations.
[0083] Figure 4 represents an example computing device 400, hereinafter referred to interchangeably as the computing system 400, in which one or more such computing devices 400 can be used to perform the process 300 of Figure 3. One or more components of the example computing device 400 can also be used to implement the system 100. The following description of the computing device 400 is provided by way of example only and is not intended to limit its application.
[0084] As shown in [Fig. 4], the example computer device 400 includes a processor 407 for executing software routines. Although only one processor is shown for clarity, the computer device 400 may also include a multiprocessor system. The processor 407 is connected to a communication infrastructure 406 for communication with other components of the computer device 400. The communication infrastructure 406 may include, for example, a communication bus, a crossbar, or a network.
[0085] The computer device 400 further includes a main memory 408, such as random access memory (RAM), and an auxiliary memory 410. The auxiliary memory 410 may include, for example, a storage drive 412, which may be a hard disk drive, a static drive, or a hybrid drive, and / or a removable storage drive 417, which may include a magnetic tape drive, an optical disc drive, a static storage drive (such as a USB flash drive, a flash memory device, a static drive, or a memory card), or the like. The removable storage drive 417 reads from and / or writes to a removable storage medium 477 in a well-known manner. The removable storage medium 477 may include a magnetic tape, an optical disc, or a memory storage medium. persistent, or similar, which is read by and written to by the removable storage drive 417. As a person skilled in the art will understand, the removable storage medium 477 includes a computer-readable storage medium in which computer-executable program instructions and / or code data are stored.
[0086] In another embodiment, the auxiliary memory 410 may include, in a complementary or alternative manner, other similar means for enabling the loading of computer programs or other instructions into the computing device 400. These means may include, for example, a removable storage unit 422 and an interface 450. Examples of removable storage units 422 and interfaces 450 include a program cartridge and cartridge interface (such as those found in video game console devices), a removable memory chip (such as an EPROM or PROM) and its associated port, a removable static storage drive (such as a USB key, flash memory device, static drive, or memory card), and other removable storage units 422 and interfaces 450 that enable the transfer of software and data from the removable storage unit 422 to the computing system 400.
[0087] The computer device 400 also includes at least one communication interface 427. The communication interface 427 allows the transfer of software and data between the computer device 400 and external devices via a communication channel 426. In various embodiments of the inventions, the communication interface 427 allows the transfer of data between the computer device 400 and a data communication network, such as a public or private data communication network. The communication interface 427 can be used to exchange data between different computer devices 400, which computer devices 400 are part of an interconnected computer network.Examples of a 427 communication interface might include a modem, a network interface (such as an Ethernet card), a communication port (such as a serial port, parallel port, printer access, GPIB, IEEE 1394, RJ45, USB), an antenna with associated circuitry, and the like. The 427 communication interface can be wired or wireless. Software and data transferred via the 427 communication interface are in the form of signals that can be electronic, electromagnetic, optical, or other signals capable of being received by the 427 communication interface. These signals are provided to the communication interface via the 426 communication channel.
[0088] As shown in [Fig.4], the computer device 400 further includes a display interface 402 which performs image rendering operations to an associated display 450 and an audio interface 452 which performs audio content playback operations via an associated speaker(s) 457.
[0089] As used herein, the term "computer program product" may refer, in part, to a removable storage medium 477, a removable storage unit 422, a hard disk installed in a storage drive 412, or a carrier wave carrying software via the communication channel 426 (wireless or cable link) to the communication interface 427. The computer-readable storage medium refers to any non-transient, persistent tangible storage medium that provides instructions and / or recorded data to the computer device 400 for execution and / or processing.Examples of such storage media include, for example, magnetic tape, CD-ROM, DVD, Blu-ray™ disc, hard disk drive, ROM or integrated circuit, static storage device (such as USB key, flash memory device, static drive or memory card), hybrid drive, magneto-optical disc, or computer-readable card such as PCMCIA card, and the like, whether such devices are internal or external to the 400 computing device.Examples of transient or intangible computer-readable transmission media that may also participate in the delivery of software, application programs, instructions and / or data to the computing device 400 include radio or infrared transmission channels as well as a network connection to another computer or networked device, and the Internet or Intranets, including email transmissions and information stored on websites, and the like.
[0090] Computer programs (also called computer program code) are stored in main memory 408 and / or auxiliary memory 410. Computer programs can also be received via the communication interface 427. When executed, these computer programs enable the computer device 400 to perform one or more features of the embodiments presented herein. In various embodiments, the computer programs, when executed, enable the processor 407 to perform the features of the embodiments described above. Therefore, these computer programs act as controllers of the computer system 400.
[0091] Software can be stored in a computer program product and loaded into the computer device 400 using the removable storage drive 417, the storage drive 412, or the interface 450. The computer program product can be a non-transient, computer-readable medium. Alternatively, the computer program product can be downloaded into the computer system 400 via the communication channel 426. When executed by the processor 407, the software causes the computer device 400 to perform the operations necessary to execute the process 300 as shown in [Fig. 3].
[0092] It is understood that the embodiment of [Fig. 4] is presented solely by way of example to explain the operation and structure of the system 400. It is therefore possible that in some embodiments, one or more features of the computer device 400 may be omitted. Similarly, in some embodiments, one or more features of the computer device 400 may be combined. Furthermore, in some embodiments, one or more features of the computer device 400 may be divided into one or more component elements.
[0093] It will be understood that the elements illustrated in [Fig.4] serve to provide means to carry out the various functions and operations of the system as described in the embodiments above.
[0094] When the computer device 400 is configured to implement the system 100 for processing one or more natural language video analytics commands, the system 100 will present a non-transient, computer-readable medium on which is stored an application which, when executed, causes the system 100 to perform steps comprising: (i) receiving, by a processing device, one or more natural language video analytics commands associated with video data, (ii) determining, using the processing device, the validity of the one or more natural language commands by means of a trained neural network, in response to a positive determination of the validity of the one or more natural language commands, (iii) generating, using the processing device,(iv) machine-readable video analytics instructions based on one or more natural language commands using the trained neural network and (iv) transmission, using the processing device, of the machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.
[0095] Those skilled in the art will understand that numerous variations and / or modifications can be made to the present invention as shown in the specific embodiments without departing from the spirit or scope of the invention as described overall. The present embodiments should therefore be considered in all respects as illustrative and not restrictive.
Claims
Demands
1. Natural language video analytics system, the system comprising: -at least one processor; and -at least one memory including computer program code; -in which the at least one processor, at least one memory and the computer program code are configured to enable the system to: -receive one or more natural language video analytics commands associated with video data; -determine the validity of the one or more natural language commands using a trained neural network; -in response to a positive determination of the validity of the one or more natural language commands, -generate machine-readable video analytics instructions based on the one or more natural language commands using the trained neural network;and -transmit machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.;
2. A system according to claim 1, wherein, in order to determine the validity of one or more natural language commands, the system is configured to: -determine the relevance of one or more natural language video analytics commands for video data analytics using the trained neural network; and -optionally, determine whether one or more natural language video analytics commands fall within the processing capabilities of the natural language video analytics system using the trained neural network.
3. System according to claim 1 or 2, wherein the machine-readable video analytics instructions include video overlay instructions, and wherein the system is configured to: -receive video data and video overlay instructions; -modify the video data based on the video overlay instructions; and -transmit the modified video data based on the video overlay instructions to the output display.
4. A system according to any one of claims 1 to 3, wherein the machine-readable video analytics instructions include video transformation instructions, and wherein the system is configured to: -receive the video data and the video transformation instructions; -modify the video data based on the video transformation instructions; and -transmit the video data modified based on the video transformation instructions to the output display.
5. A system according to any one of claims 1 to 4, wherein the system is further configured to: - compare machine-readable video analytics instructions to an access control list; - in response to a comparison result indicating authorized access, - retrieve video-derived data from one or more databases associated with the video dataset, based on machine-readable video analytics instructions; - generate one or more data representations using the retrieved data and machine-readable video analytics instructions; and - transmit the one or more data representations to the display module.
6. System according to claim 5, wherein the system is further configured to: -generate a text response based on the retrieved data and machine-readable video analytics instructions; and -transmit the text response to the display module.
7. A system according to claim 5, wherein the system is further configured to: - generate video-derived data based on a predetermined set of video analytics instructions using video data and a trained video analytics algorithm; and -store the video-derived data and the video data in one or more databases.
8. Method for processing one or more natural language video analytics commands, the method comprising: -receiving, by a processing device, one or more natural language video analytics commands associated with video data; -determining, using the processing device, the validity of one or more natural language commands with the help of a trained neural network; -in response to a positive determination of the validity of one or more natural language commands, -generating, using the processing device, machine-readable video analytics instructions based on the one or more natural language commands with the help of the trained neural network;and -the transmission, using the processing device, of machine-readable video analytics instructions to a display module, the display module being configured to update an output display based on the machine-readable video analytics instructions.;
9. A method according to claim 8, wherein the step of determining the validity of one or more natural language commands using the trained neural network comprises one or more of the steps of: -determining, using the processing device, the relevance of one or more natural language video analytics commands for video data analytics using the trained neural network; and -determining, using the processing device, whether one or more natural language video analytics commands fall within a processing capability of the natural language video analytics system using the trained neural network.
10. A method according to claim 8 or 9, wherein the machine-readable video analytics instructions include video overlay instructions, and wherein the method further comprises: - the reception, by the display module, of the video data and the video overlay instructions, - the modification, using the display module, of video data based on video overlay instructions; and - the transmission, using the display module, of the video data modified based on video overlay instructions to the output display.
11. A method according to any one of claims 8 to 10, wherein the machine-readable video analytics instructions include video transformation instructions, and wherein the method further comprises: - the receipt, by the display module, of the video data and the video transformation instructions, - the modification, using the display module, of the video data on the basis of the video transformation instructions; and - the transmission, using the display module, of the video data modified on the basis of the video transformation instructions to the output display.
12. A method according to any one of claims 8 to 11, further comprising the steps of: - comparing, using the processing device, machine-readable video analytics instructions with an access control list; - in response to a comparison result indicating authorized access, - retrieving, using the processing device, video-derived data from one or more databases associated with the video data, based on the machine-readable video analytics instructions;- generation, using the processing device, of one or more data representations using the retrieved data and machine-readable video analytics instructions, and - transmission, using the processing device, of one or more data representations to the display module, the display module being configured to update the output display on the basis of one or more data representations.;
13. A method according to claim 12, further comprising the steps of: - generation, using the processing device, of a textual response based on the retrieved data and machine-readable video analytics instructions; and - transmission, using the processing device, of the textual response to the display module.
14. A method according to claim 12, further comprising the steps of: -generating, using the processing device, video-derived data based on a predetermined set of video analytics instructions using the video data and a trained video analytics algorithm; and -storing the video-derived data and the video data in one or more databases.