Systems and Methods for Distributed Data Analysis

The method and system for generating device-specific ANN models on user devices address the challenge of efficiently processing large datasets by optimizing neural network distribution and interaction, enhancing data analysis efficiency and accuracy across diverse devices.

JP7701379B2Active Publication Date: 2025-07-01ゼイリエント
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022567287
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-08
Filing Date
2021-05-10
Publication Date
2025-07-01
Estimated Expiration
2041-05-10

AI Technical Summary

Technical Problem

Existing data analysis platforms face challenges in efficiently processing large and diverse datasets using neural networks due to computational intensity and difficulty in distributing complex data analysis tools across multiple devices.

Method used

A method and system for generating device-specific artificial neural network (ANN) models that can be distributed across user devices like smartphones and IoT devices, utilizing a unified platform for data analysis, including user interfaces, data storage, and AI/ML modules to identify objects of interest and optimize model execution.

Benefits of technology

Enhances data processing efficiency, improves model accuracy, and facilitates wider adoption of data analysis tools by enabling easier interaction and distribution of ANN models tailored to specific device characteristics and use cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701379000001
    Figure 0007701379000001
  • Figure 0007701379000002
    Figure 0007701379000002
  • Figure 0007701379000003
    Figure 0007701379000003
Patent Text Reader

Abstract

The present invention provides a system and method for generating device-specific artificial neural network (ANN) models for distribution across user devices (105). A sample dataset (140) is collected from devices in a particular environment or use case and includes predictions from device-specific ANN models running on the user devices. The received dataset is used in conjunction with existing datasets and stored ANN models to generate updated device-specific ANN models from each stored instance of the device ANN model based on the training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 021,735, filed on May 8, 2020, entitled “Systems and Methods for Distributed Data Analytics,” the entire disclosure of which is incorporated herein by reference.

[0002] The following disclosure is directed to methods and systems in data analytics, and more specifically, to the distribution of data analysis frameworks and related data tools.

Background Art

[0003] With the development of intelligent systems, the amount of data that is read, transmitted, and further processed is continuously increasing.

[0004] Complex data analysis can implement a machine learning mechanism and use a large training data set to train a neural network. These neural networks can then be used to process input data within a similar domain to the large training data set. Applying such complex neural network logic to a broader and larger data set of input data can prove to be difficult and computationally intensive. Thus, as disclosed herein, methods and systems for providing access to data processing tools and a distributable data analysis platform provide such systems with the benefits of higher platform adoption and distribution rates, greater data availability, improved training effectiveness, and improved execution efficiency.

[0005] Accordingly, improved methods and systems for data processing using neural networks can significantly benefit from improved execution efficiency.

Summary of the Invention

Means for Solving the Problem

[0006] Current data analysis platforms use various external tools to perform specific tasks. This disclosure describes techniques and related systems that facilitate complex data analysis using tools that can be remotely accessed or otherwise distributed across multiple devices in a uniform platform using a defined analysis framework.

[0007] What is disclosed herein are exemplary embodiments of methods and support systems that facilitate data analysis using a wide range of data storage devices, neural networks, and other data science tools. Providing access to such powerful tools supports a wider adoption base and thus results in larger and more accurate datasets for training and development purposes. The ease of user interaction depends primarily on the user interface provided by the platform and the methods used to provide access to powerful data analysis tools, model training mechanisms, and neural networks implemented therein. The various systems and methods provided by the present invention operate in conjunction to process input data, identify regions of interest and / or objects of interest within the regions, while providing the user with an easy-to-view and easy-to-read interface to scrutinize input and training data, improve the accuracy of the models used by the network, and visualize the results, actively employing a number of neural networks.

[0008] The object can be an inanimate object that is generally identified (e.g., "car" or "pedestrian") or a specific object that can be specifically identified, based on, for example, a combination of face recognition, character recognition, or similar techniques.

[0009] Accordingly, on a first aspect, the present invention provides a method for generating a device-specific artificial neural network (ANN) model for distribution across user devices such as smartphones, cameras, and other Internet of Things (IoT) devices. In various embodiments, the method includes steps of a processor receiving a sample dataset from a user device in a user environment, the sample dataset comprising media data and a prediction by a device-specific ANN model executed on the user device, and the processor writing the sample dataset to a training data storage device. The method also includes steps of the processor identifying, within the data storage device, (i) a use case dataset comprising at least training data parameters, (ii) training data from the sample dataset that satisfies the training data parameters provided within the use case dataset, and (iii) a stored instance of the device-specific ANN model. The processor then generates, at least in part, an updated device-specific ANN model from each of the stored instances of the device ANN model based on the identified training data. In some cases, a library of device-specific parameters and training data is maintained and the generation of the updated device-specific ANN model is further based on the device-specific parameters and training data.

[0010] In some embodiments, the media data comprises image data and the application of the ANN model to the image data facilitates the identification of objects of interest within the image data. The training data parameters may include media data parameters such as color index, brightness index, contrast index, image temperature, color tone, one or more hue values, and / or gamma value, and / or device parameters such as available memory, processing speed, image resolution, and / or capture frame rate.

[0011] In some cases, the use case dataset is specific to a particular use case and, in some instances, includes environmental aspects (such as installation of the device in an outdoor environment, installation of the device in an indoor environment, installation of the device in a well-lit environment, or installation of the device in a poorly lit environment, etc.) and functional aspects (such as, for example, face recognition, character recognition, document certification, etc.). In some embodiments, the predictions generated by the device-specific ANN model include a per-image quantitative image saliency metric indicating the likelihood that the media file contains the object of interest, and in some cases, at least partially based on the per-image quantitative image saliency metric, determine the minimum number of images necessary to achieve a threshold model accuracy.

[0012] In some cases, the method further includes distributing the device-specific updated ANN model to at least a subset of the user devices associated therewith.

[0013] In another aspect, the present invention provides a system for generating a device-specific artificial neural network (ANN) model for distribution across user devices such as smartphones, cameras, and other Internet of Things (IoT) devices. The system includes one or more processors and a memory coupled to the processors, and the processors execute a plurality of modules stored in the memory. The modules include a user interface that receives instructions from a user, the instructions identifying one or more sample data sets from user devices in the user environment, the sample data sets comprising media data and predictions by a device-specific ANN model executed on the user devices; a data storage device comprising the sample data sets; a business logic module that, when executed, (i) identifies a use case data set stored in the data storage device, the use case data set comprising at least training data parameters, (ii) identifies training data from the sample data sets that meet the training data parameters provided in the use case data set, and (iii) identifies a device-specific ANN model stored in the data storage device; and an artificial intelligence machine learning module that, when executed, generates an updated device-specific ANN model from each of the stored instances of the device ANN model based on the training data.

[0014] In some embodiments, the media data comprises image data, and the application of the ANN model to the image data facilitates the identification of objects of interest within the image data. The training data parameters may include media data parameters such as color index, brightness index, contrast index, image temperature, color tone, one or more hue values, and / or gamma value, and / or device parameters such as available memory, processing speed, image resolution, and / or capture frame rate.

[0015] In some cases, the use case dataset is specific to a particular use case and, in some instances, includes environmental aspects (such as installation of the device in an outdoor environment, installation of the device in an indoor environment, installation of the device in a well-lit environment, or installation of the device in a poorly-lit environment, etc.) and functional aspects (such as, for example, face recognition, character recognition, document certification, etc.). In some embodiments, the predictions generated by the device-specific ANN model include a per-image quantitative image saliency metric that indicates the likelihood that the media file contains the object of interest, and in some cases, at least partially based on the per-image quantitative image saliency metric, determine the minimum number of images necessary to achieve a threshold model accuracy.

[0016] In some cases, the distribution module distributes the device-specific updated ANN model to at least a subset of the user devices associated therewith.

[0017] In another aspect, the present invention provides a method for optimizing the execution of a device-specific trained artificial neural network (ANN) model on an edge device (such as a smartphone, camera, and other Internet of Things (IoT) devices, etc.), which includes steps of receiving, by a processor, a first trained ANN model and a second ANN model, wherein the first ANN model and the second ANN model each perform different estimations on input data, and the output of the first ANN model serves as an input to the second ANN model, and merging, according to control flow instructions, the first ANN model, the second ANN model, and the control flow execution instructions into a combined software package for deployment to the edge device for execution thereon.

[0018] In one embodiment, the first trained ANN model and the second trained ANN model each comprise an individual analysis criterion and use case data, and the processor selects the first and second ANN models at least in part based on the analysis criteria therein. The parent ANN may be generated as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and the meta-architecture can then be delivered to the edge device for execution as a single ANN model. In an embodiment, the edge device is a camera, and execution of the first ANN model and the second ANN model on the camera can identify the object of interest in the image file captured on the camera.

[0019] In another aspect, the present invention provides a system for optimizing the execution of device-specific trained artificial neural network (ANN) models on edge devices (such as smartphones, cameras, and other Internet of Things (IoT) devices). The system includes one or more processors and a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory. When executed, the instructions identify a first trained ANN model and a second trained ANN model in a data storage device, the first ANN model and the second ANN model each performing different estimations on input data, the output of the first ANN model serving as an input to the second ANN model, merging the first ANN model, the second ANN model, and control flow execution instructions into a combined software package, and using a distribution module to deploy the combined software package to the edge device for execution thereon according to the control flow instructions.

[0020] In one embodiment, the first trained ANN model and the second trained ANN model each comprise an individual analysis criterion and use case data, and the processor selects the first and second ANN models at least in part based on the analysis criteria therein. The parent ANN may be generated as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and the meta-architecture can then be delivered to the edge device so that it can be executed as a single ANN model. In an embodiment, the edge device is a camera, and the execution of the first ANN model and the second ANN model on the camera can identify the object of interest in the image file captured on the camera.

[0021] In another aspect, the present invention provides a method for identifying an object of interest in an image file. The method includes receiving one or more image files, each image file potentially including an object of interest, and applying a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values indicating the likelihood that a particular pixel is part of the object of interest. Based on the ground truth label, a three-dimensional saliency surface map having an x-axis, a y-axis, and a z-axis is generated, where the x-axis and y-axis values define the location of the pixels in the image and the z-axis value is the pixel-specific saliency value. A curve shape is selected from a library of curve shapes, the curve shape is applied to the saliency surface map, the fit between the curve shape and the three-dimensional surface is determined, and based on the fit, it is determined whether the image file includes an object of interest.

[0022] In some embodiments, the curve shape is selected based on the object of interest and may be at least partially based on one or more statistical distributions such as a Gaussian distribution, a Poisson distribution, or a hybrid distribution. In some cases, an image file is added to a library of image files for use in training an artificial neural network (ANN), and the ANN is trained to identify the object of interest in subsequent media files and / or segment objects in subsequent media files.

[0023] In another aspect, the present invention provides a system for identifying an object of interest in an image file, the system including one or more processors and a memory coupled to the one or more processors, the one or more processors configured to execute computer-executable instructions stored in the memory. When executed, the system receives one or more image files, each image file potentially including an object of interest, applies a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values indicating the likelihood that a particular pixel is part of the object of interest. Based on the ground truth label, a three-dimensional surface having an x-axis, a y-axis, and a z-axis is generated, the x-axis and y-axis values defining the location of pixels in the image and the z-axis value being the pixel-specific saliency value. A curve shape is selected from a library of curve shapes, the curve shape is applied to the ground truth label, a fit between the curve shape and the three-dimensional surface is determined, and based on the fit, it is determined whether the image file includes an object of interest.

[0024] In some embodiments, the curve shape is selected based on the object of interest and may be at least partially based on one or more statistical distributions such as a Gaussian distribution, a Poisson distribution, or a hybrid distribution. In some cases, an image file is added to a library of image files for use in training an artificial neural network (ANN), and the ANN is trained to identify the object of interest in subsequent media files and / or segment the objects in subsequent media files.

[0025] In yet another aspect, the present invention provides a method for storing image data for the transmission of video data, including receiving video data in a standard video data format (such as H.264) at an edge device, and extracting an image slice from the video data, the image slice comprising an image, a start index time and an end index time indicating the temporal location of the image slice in the video data, and a region-of-interest parameter describing the two-dimensional coordinates of the region of interest in the image.

[0026] In some embodiments, the receiving of video data and the extraction of image slices are performed on the edge device. The image slices may then be analyzed using one or more artificial neural networks on the edge device to determine the region of interest and whether the region of interest contains the object of interest. In some cases, the image slices are identified as high resolution if the image slice contains the object of interest and low resolution otherwise. The method may further include transmitting the high-resolution image slices to an artificial intelligence machine learning module for inclusion in a training data set specific to the edge device on which the image was captured.

[0027] In another aspect, the present invention provides a system for storing image data for the transmission of video data, comprising one or more processors and a memory coupled to the one or more processors, wherein the one or more processors execute computer-executable instructions stored in the memory. When the instructions are executed, the system receives video data in one of a plurality of standard video data formats (e.g., H.264) at an edge device, extracts image slices from the video data, and the image slices comprise an image, a start index time and an end index time indicating the temporal location of the image slice in the video data, and region-of-interest parameters describing the two-dimensional coordinates of the region of interest in the image.

[0028] In certain embodiments, the receiving of video data and the extraction of image slices are performed on the edge device. The image slices may then be analyzed using one or more artificial neural networks on the edge device to determine the region of interest and whether the region of interest contains the object of interest. In some cases, the image slice is identified as high resolution if the image slice contains the object of interest and low resolution otherwise. The method may further include transmitting the high-resolution image slices to an artificial intelligence machine learning module for inclusion in a training data set specific to the edge device on which the image was captured.

[0029] Features described in the context of separate aspects and / or embodiments of the present invention may, where possible, be used together and / or be interchangeable with each other. Similarly, where features are described in the context of a single embodiment for the sake of brevity, these features may also be provided separately or in any suitable sub-combination. Features described in connection with a system may have corresponding features definable and / or combinable in connection with a method, or vice versa, and these embodiments are specifically contemplated. The present invention provides, for example, the following. (Item 1) A method for generating a device-specific artificial neural network (ANN) model for distribution across user devices, the method comprising: receiving, by a processor, a sample data set from the user device in the user environment, the sample data set comprising media data and a prediction by a device-specific ANN model executed on the user device; writing, by the processor, the sample data set to a training data storage device; identifying, by the processor, within the data storage device, a use case data set, the use case data set comprising at least training data parameters; identifying, by the processor, within the training data storage device, training data from the sample data set that satisfies the training data parameters provided within the use case data set; identifying, by the processor, within the data storage device, a stored instance of the device-specific ANN model; generating, by the processor, an updated device-specific ANN model from each of the stored instances of the device ANN model based on the training data A method comprising. (Item 2) The method according to item 1, wherein the user device comprises a plurality of image capture devices. (Item 3) The method according to item 1 or item 2, wherein the media data comprises image data, and application of the ANN model to the image data facilitates identification of an object of interest within the image data. (Item 4) The method according to item 1, item 2, or item 3, wherein the training data parameters include media data parameters and device parameters. (Item 5) The method according to item 4, wherein the media data parameters include one or more of a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and a gamma value. (Item 6) The method according to item 4 or item 5, wherein the device parameters include one or more of available memory, processing speed, image resolution, and capture frame rate. (Item 7) The use case dataset is the method described in either Item 1 or Items 2-6, which is specific to a particular use case. (Item 8) The use case is the method described in Item 7, which includes an environmental aspect and a functional aspect. (Item 9) The functional aspect of the use case is the method described in Item 8, which includes face recognition. (Item 10) The environmental aspect of the use case is the method described in Item 8 or Item 9, which includes one of the following: installation of the device in an outdoor environment, installation of the device in an indoor environment, installation of the device in a well-lit environment, or installation of the device in a poorly-lit environment. (Item 11) The prediction generated by the device-specific ANN model is the method described in either Item 1 or Items 2-10, which includes a per-image quantitative image saliency metric indicating the likelihood that the media file contains the object of interest. (Item 12) The method described in Item 11 further includes determining, at least in part, the minimum number of images required to achieve a threshold model accuracy based on the per-image quantitative image saliency metric. (Item 13) The method described in either Item 1 or Items 2-12, which includes maintaining a library of device-specific parameters and training data, and generating the updated device-specific ANN model further based on the device-specific parameters and training data. (Item 14) The method described in either Item 1 or Items 2-13 further includes distributing, by the processor, the updated device-specific ANN model to at least a subset of the user devices associated therewith. (Item 15) A system for generating a device-specific artificial neural network (ANN) model for distribution across user devices, the system comprising: one or more processors; a memory coupled to the one or more processors, wherein the one or more processors execute a plurality of modules stored in the memory, and the plurality of modules comprise: A user interface that receives commands from a user, the command identifying one or more sample data sets from the user device in the user environment, the sample data set comprising media data and a prediction by a device-specific ANN model executed on the user device, the user interface; A data storage device comprising the sample data set; A business logic module that, when executed, (i) identifies a use case data set stored in the data storage device, the use case data set comprising at least training data parameters, (ii) identifies training data from the sample data set that satisfies the training data parameters provided in the use case data set, and (iii) identifies a device-specific ANN model stored in the data storage device, the business logic module; An artificial intelligence machine learning module that, when executed, generates an updated device-specific ANN model from each of the stored instances of the device ANN model based on the training data, the artificial intelligence machine learning module Comprising a memory Comprising a system. (Item 16) The system according to item 15, wherein the user device comprises a plurality of image capture devices. (Item 17) The system according to item 16, wherein the media data comprises image data, and the application of the ANN model to the image data facilitates the identification of objects of interest in the image data. (Item 18) The system according to item 15, item 16, or item 17, wherein the training data parameters include media data parameters and device parameters. (Item 19) The system according to item 18, wherein the media data parameters include one or more of a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and a gamma value. (Item 20) The system according to item 18 or item 19, wherein the device parameters include one or more of available memory, processing speed, image resolution, and capture frame rate. (Item 21) The use case dataset is the system described in either item 15 or items 16 - 20, which is specific to a particular use case. (Item 22) The use case is the system described in item 21, which includes an environmental aspect and a functional aspect. (Item 23) The functional aspect of the use case is the system described in item 22, which includes face recognition. (Item 24) The environmental aspect of the use case is the system described in item 22 or item 23, which includes one of the following: installation of the device in an outdoor environment, installation of the device in an indoor environment, installation of the device in a well - lit environment, or installation of the device in a poorly - lit environment. (Item 25) The prediction generated by the device - specific ANN model is the system described in either item 15 or items 16 - 24, which includes a per - image quantitative image saliency metric indicating the likelihood that the media file contains the object of interest. (Item 26) The artificial intelligence machine learning module further determines, at least in part, the minimum number of images required to achieve a threshold model accuracy based on the per - image quantitative image saliency metric, as described in item 25. (Item 27) The system further includes a library of device - specific parameters and training data, and the artificial intelligence machine learning module generates the updated device - specific ANN model based on the device - specific parameters and training data, as described in either item 15 or items 16 - 26. (Item 28) The system further includes a deployment module for distributing the updated device - specific ANN model to at least a subset of the user devices associated therewith, as described in either item 15 or items 16 - 27. (Item 29) A method for optimizing the execution of a device - specific trained artificial neural network (ANN) model on an edge device, the method comprising: Receiving, by a processor, a first trained ANN model and a second ANN model, wherein the first ANN model and the second ANN model each perform different estimations on input data, and the output of the first ANN model serves as an input to the second ANN model. merging into a software package combined with the first ANN model, the second ANN model, and control flow execution instructions; deploying the combined software package to an edge device for execution thereon according to the control flow instructions; A method comprising. (Item 30) The method according to item 29, wherein the first trained ANN model and the second trained ANN model each comprise an individual analysis criterion and use case data, and the processor selects the first and second ANN models at least in part based on the analysis criterion therein. (Item 31) The method according to item 29 or item 30, further comprising generating a parent ANN as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, wherein the meta-architecture is delivered to the edge device so that it executes as a single ANN model. (Item 32) The method according to item 29, item 30, or item 31, wherein the edge device comprises a camera. (Item 33) The method according to item 32, wherein execution of the first ANN model and the second ANN model on the camera identifies an object of interest in an image file captured on the camera. (Item 34) A system for optimizing the execution of device-specific trained artificial neural network (ANN) models on an edge device, the system comprising: one or more processors; a memory coupled to the one or more processors, wherein the one or more processors execute computer-executable instructions stored in the memory, and the computer-executable instructions, when executed, identifying in a data storage device a first trained ANN model and a second ANN model, wherein the first ANN model and the second ANN model each perform different estimations on input data, and an output of the first ANN model serves as an input to the second ANN model; merging into a software package combined with the first ANN model, the second ANN model, and control flow execution instructions; The dispersion module, in accordance with the control flow instruction, for execution thereon, deploy the combined software package to the edge device causing it to be done, memory and A system comprising (Item 35) The first trained ANN model and the second trained ANN model each comprise an individual analysis criterion and use case data, and the processor selects the first and second ANN models at least in part based on the analysis criterion therein, the system according to item 34 (Item 36) Execution of the instruction further generates a parent ANN as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and the meta-architecture is delivered to the edge device for execution as a single ANN model, the system according to item 34 or item 35 (Item 37) The edge device comprises a camera, the system according to item 34, item 35, or item 36 (Item 38) Execution of the first ANN model and the second ANN model on the camera identifies the object of interest in the image file captured on the camera, the system according to item 37 (Item 39) A method for identifying an object of interest in an image file, the method comprising Receiving one or more image files, each image file potentially containing an object of interest Applying a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values indicating the likelihood that a particular pixel is part of the object of interest Generating a three-dimensional saliency surface map having an x-axis, a y-axis, and a z-axis, where the x-axis and y-axis values define the location of pixels in the image and the z-axis value is the pixel-specific saliency value Selecting a curve shape from a library of curve shapes, applying the curve shape to the saliency surface map, and determining the fit between the curve shape and the three-dimensional surface Based on the fit, determining whether the image file contains the object of interest Including, method (Item 40) The method according to item 39, wherein the curve shape is selected based on the object of interest. (Item 41) The method according to item 39 or item 40, wherein the curve shape is selected from one of a Gaussian distribution, a Poisson distribution, and a hybrid distribution. (Item 42) The method according to item 39, item 40, or item 41, further comprising adding the image file to a library of image files for use in training an artificial neural network (ANN). (Item 43) The method according to item 40, item 41, or item 42, wherein the ANN is trained to identify an object of interest in a subsequent media file. (Item 44) The method according to item 40 or any one of items 41 - 43, wherein the ANN is trained to segment objects in a subsequent media file. (Item 45) A system for identifying an object of interest in an image file, the system comprising: one or more processors; a memory coupled to the one or more processors, wherein the one or more processors execute computer - executable instructions stored in the memory, and the computer - executable instructions, when executed, receive one or more image files, each image file potentially including an object of interest; apply a non - binary ground - truth label to each image file, the non - binary ground - truth label comprising a distribution of pixel - specific saliency values indicating the likelihood that a particular pixel is part of the object of interest; generate a three - dimensional saliency surface map having an x - axis, a y - axis, and a z - axis, wherein the x - axis and y - axis values define the location of pixels in the image and the z - axis value is the pixel - specific saliency value; select a curve shape from a library of curve shapes, apply the curve shape to the saliency surface map, and determine a fit between the curve shape and the three - dimensional surface; determine whether the image file contains the object of interest based on the fit; and a memory for causing the above to be performed. A system comprising the above. (Item 46) The system according to item 45, wherein the curve shape is selected based on the object of interest. (Item 47) The system according to item 45 or item 46, wherein the curve shape is selected from one of a Gaussian distribution, a Poisson distribution, and a hybrid distribution. (Item 48) The system according to item 45, item 46, or item 47, wherein the execution of the instruction further adds the image file to a library of image files for use in training an artificial neural network (ANN). (Item 49) The system according to item 48, wherein the ANN is trained to identify an object of interest in a subsequent media file. (Item 50) The system according to item 48 or item 49, wherein the ANN is trained to segment an object in a subsequent media file. (Item 51) A method for storing image data for transmission of video data, the method comprising: receiving video data in one of a plurality of standard video data formats at an edge device; extracting a plurality of image slices from the video data, the image slices comprising an image, a start index time and an end index time indicating the temporal location of the image slice in the video data, and region-of-interest parameters describing the two-dimensional coordinates of a region of interest in the image; A method comprising the above. (Item 52) The method according to item 51, wherein the receiving of the video data and the extraction of the image slices are performed on an edge device. (Item 53) The method according to item 52, further comprising using one or more artificial neural networks on the edge device to analyze the image slices and determine the region of interest and whether the region of interest contains an object of interest. (Item 54) The method according to item 53, further comprising identifying each image slice as high resolution if the image slice contains an object of interest, and as low resolution otherwise. (Item 55) The method according to item 54, further comprising transmitting the high-resolution image slices to an artificial intelligence machine learning module for inclusion in an artificial neural network of a training data set specific to the edge device on which the image was captured. (Item 56) The method according to item 51 or any of items 52-55, wherein the standard video data format comprises an H.264 data format. (Item 57) A system for storing image data for the transmission of video data, the method comprising: one or more processors; a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory, the computer-executable instructions, when executed, receiving video data at an edge device in one of a plurality of standard video data formats; extracting a plurality of image slices from the video data, the image slices comprising an image, a start index time and an end index time indicating the temporal location of the image slice within the video data, and region-of-interest parameters describing the two-dimensional coordinates of the region of interest within the image; causing a memory to comprise a system. (Item 58) The system of item 57, wherein the receiving of the video data and the extraction of the image slices are performed on an edge device. (Item 59) The system of item 58, wherein the execution of the computer-executable instructions further uses one or more artificial neural networks on the edge device to analyze the image slices and determine the region of interest and whether the region of interest contains a region-of-interest object. (Item 60) The system of item 59, wherein the execution of the computer-executable instructions further identifies each image slice as high resolution if the image slice contains a region-of-interest object and as low resolution otherwise. (Item 61) The system of item 60, wherein the execution of the computer-executable instructions further transmits the high-resolution image slices to an artificial intelligence machine learning module for inclusion in an artificial neural network of a training data set specific to the edge device on which the image was captured.

Brief Description of the Drawings

[0030] In the drawings, like reference numerals generally refer to the same parts throughout different figures. Also, the drawings are not necessarily to scale; instead, emphasis is generally placed on illustrating the principles of the implementation. In the following description, various implementations are described with reference to the following drawings.

[0031]

Figure 1

[0032]

Figure 2

[0033]

Figure 3

[0034]

Figure 4

DETAILED DESCRIPTION OF THE INVENTION

[0035] Detailed Description What is described herein is, in one embodiment, a method and support system for generating, deploying, and further maintaining an endpoint-deployable artificial intelligence system, a machine learning mechanism, and a data model implemented as an inclusive platform. As shown in FIG. 1, the platform 100 implements a framework that contains a front-end user interface ("user interface") 105 for interaction with a user, a business logic module 110, a data storage device 115, an artificial intelligence / machine learning ("AI / ML") training module 120, and a deployment tool integrated within a user environment 125. The framework components may be communicatively coupled using one or more APIs (130a, 130b, 130c, and 130d) as provided by the platform 100.

[0036] According to some embodiments, the platform 100 provides one or more user interfaces 105 for a user of the platform to access a dataset 140 provided by the user and collected from an endpoint device, and otherwise perform data analysis thereon. These user interfaces 105 may be provided, among other things, using distributed, localized applications (e.g., SDKs, APKs, IPAs, JVM files, other localized executable files, and equivalents), APIs (e.g., JSON, REST, other data transfer protocols, and equivalents), websites, or web application functionality, either in combination or separately. The user interface 105 facilitates the provision of analysis criteria to a reference collection system. The analysis criteria may include configuration, parameters, and access to the user's dataset. According to some embodiments, the configuration and parameters may be used as, or otherwise referred to as, use case data. Use cases may include functional processes such as face recognition, license plate and other character recognition, image detection processes for identity document authentication, object detection for autonomous driving applications, motion detection and intruder alerts, and others. Use cases may also include environmental aspects such as outdoor to indoor installation, night to day, dense spaces (e.g., airports, transfer stations) to sparse spaces (e.g., security cameras for banks, home cameras), and others.

[0037] Importantly, the edge devices used within each use case may vary and, in many cases in many embodiments of the present invention, have device-specific characteristics and processing limitations that are considered by, and / or incorporated into, the models used in those devices. Examples of device-specific characteristics can include device-specific properties such as available memory, processing speed, image resolution, capture frame rate, and others.

[0038] For example, the user interface 105 can provide the user with analysis feedback based on the uploaded data set 140. The data set provided by the user (referred to herein as the "media data set") may include, but is not limited to, among other things, a single image file, multiple image files, a composite image file containing multiple images therein (e.g., among other things, GIF, APNG, WebP), a video file containing one or more frames, multiple video files, and audio files. The feedback may include data for characterizing the data set before performing further operations such as classifying and balancing the image data set on such a data set. The feedback may be qualitative (e.g., high quality, low quality, etc.) or quantitative, such as one or more quality metrics for the training data set, which, by nature, describe various photometric properties (brightness, luminance, color spectrum, etc.) and geometric properties (shape, edge definition, etc.) of the image and potential objects of interest within the image.

[0039] The data set containing the image may be further analyzed to extract or otherwise generate media properties of the image and other objects contained therein. The media properties may include, but are not limited to, among other things, color index, brightness index, contrast index, and other image properties (e.g., temperature, color tone, hue, gamma). A data set containing more than one image, such as a composite image file or a video file, may be analyzed as a batch to identify, extract, or otherwise generate media properties for the multiple image files or video files of the data set.

[0040] Platform 100 also generates other media properties, such as complexity indices, from media datasets provided by a user. The complexity index may be a set of diagnostic data representing the complexity of an image of a media dataset or a frame of one or more videos. The platform may further compare media properties within media datasets provided by a user, such as images, frames of video files, things associated with frames between video files, or between video files themselves. The user interface 105 of platform 100 may also be used to identify or further generate comparisons of media properties or other characteristics within or between media datasets. For example, the platform can generate a comparison of the background and foreground of an image dataset, such as that found in an individual image or a frame of an individual video. Similarly, the platform can also generate a comparison of objects of interest and other objects not of interest, such as those contained within a media dataset, such as distinguishing people within an image from background objects. Further, the platform may assign classes to media datasets for further comparison therebetween. Examples of classes may include general categories such as people, human faces, cars, animals, defects in manufactured goods, or specific classes such as people in a certain vicinity, people within a certain distance, adult German Shepherds, adult Dalmatians, young Labrador puppies, or cracks in a material, contaminated material, or chips on a material.

[0041] In other embodiments, the platform 100 can generate a quantitative image saliency metric for an image within the image dataset, which may comprise a single number or a matrix or region or other measurement assigned at the pixel level, which can be used to predict difficulty, and using which the calculation process can distinguish among the objects of interest and / or between the objects of interest and the background within the image. Based on the image saliency metric, the minimum number of images necessary to train the model to achieve a defined accuracy can be determined. The process can be extended using human-readable criteria such as brightness, contrast, distance from the camera, etc., and can provide additional image collection recommendations to further improve and refine the training dataset. For example, the platform may identify that while the training data contains a set of dark / far and dark / near images, adding brighter / far images would result in a significantly improved training dataset. Similarly, if the training data contains high-quality images with significant contrast values, adding additional images to the training data may not be necessary or may only marginally increase the accuracy of the model.

[0042] According to some embodiments, the user interface 105 can provide recommendations to the user based on feedback associated with its media dataset 140. In some examples, the recommendations may be provided with the feedback or, alternatively, may be included therein. Recommendations such as those provided by the platform may include, without limitation, proposals for additional data for the user to collect and include within the media dataset, as well as proposed extensions to be applied to the one or more media datasets for improving it.

[0043] According to some embodiments, the analysis performed by platform 100 may be implemented by a machine learning mechanism or an artificial neural network (「ANN」). To implement such analysis, the platform may further include a criteria collection system and may provide the user with access to artificial intelligence tools and the ability to use a front-end user interface. For example, one or more user interfaces may be provided to collect from the user the primary analysis criteria regarding requirements or other preferences by which the platform may use the user's media dataset for analysis. For example, the user may identify analysis criteria including, but not limited to, speed and latency requirements as required by the user's implementation, hardware and network requirements as required by the user's implementation, the size of objects to be identified within the media dataset, reaction time tolerances as required by the user's implementation, tolerance regarding false detection as identified by the platform, tolerance regarding non-detection as identified by the platform, and accuracy requirements for predictions made by the platform. In some instances, the criteria collection system may also facilitate filtering of large datasets to a dataset that meets certain image criteria or size limits.

[0044] According to some embodiments, the platform uses the platform's intelligent system (e.g., machine learning mechanisms, artificial neural networks, and equivalents) to identify the key analysis criteria that are best suited for the user's implementation. Some embodiments of the platform's criteria collection system use a dual (or multiple) ANN to provide the user with access to the best artificial intelligence tools and the capabilities for their associated use cases. In other words, the first neural network may receive a media dataset as provided by the user and determine the best analysis criteria to be used by the second neural network to perform a specific analysis on the same or other media datasets as provided by the user. For example, the user may upload a video clip of a sample use case to the first neural network. The user may identify a common use case or object, whether selected from a list or identified in a custom manner by the user, for use in video clip analysis. Based on the user's selection, the first neural network analyzes the video clip provided by the user and determines the analysis criteria necessary for the second neural network to more appropriately analyze the uploaded video clip. For example, the first ANN may be used to identify the area of interest within an image that is likely to contain a person, among a plurality of other objects in the image, while the second ANN may process the area of interest and be used to perform face recognition on the image of the person. In some cases, the analysis criteria may be automatically extracted from the video clip and may include the reaction time, a specific definition of the accuracy metric, and the quantitative value of the metric. The first neural network may determine the required "reaction time" of the second neural network, the size of the object to be identified by the second neural network, or further, the ideal number of frames of the video that the second neural network may use at runtime to correctly determine the "reaction".

[0045] According to some embodiments, the platform may further include intelligent operation tools to facilitate the implementation and maintenance of the user's intelligent system (e.g., machine learning mechanisms, artificial neural networks, and equivalents). For example, the platform may provide the user with an integrated compilation of a software application or software development kit (SDK) for the user's specific target hardware. The SDK compilation may contain a unique license (e.g., a token) built into or associated with it, however, other licensing models may also be used. The software (e.g., SDK, other software applications, etc.) facilitates the monitoring of statistical information and / or its performance regarding the software and communication between the hardware that executes the software and various platform components. The software may further provide a comparison of statistical information regarding the view of statistical information about the training data.

[0046] In some embodiments, the platform uses data obtained by software that is distributed across the user's hardware to provide the user with recommendations regarding training data and the configuration of the intelligent system. The platform may also provide a media dataset such as that collected at runtime by the user's hardware with predictions superimposed therein. By doing so, the platform may further provide a user interface for the user to mark whether the predictions provided in the runtime data are correct, incorrect, or, in some cases, rank them along a scale of accuracy (e.g., a numerical value, probability, qualitative tag, etc. representing the likelihood that the prediction is correct). In response to receiving an indication that a prediction is correct, the platform may add the associated runtime data to the auxiliary training dataset. Adding runtime data with correct or corrected predictions to the auxiliary training dataset facilitates the continued training of a semi-supervised machine learning procedure that updates the ANN model (or other artificial intelligence model) for use by the user's intelligent system. Once updated, the ANN model is deployed to the user's hardware and distributes the improvement to the user's intelligent system.

[0047] The AI / ML training system accesses a user's dataset such as provided by the user and generates a subsample of the training data according to the configuration and parameters of the analysis criteria. For example, the configuration and parameters as provided by the analysis criteria may include requirements for limiting the training data to a dataset with faces close to the camera and excluding faces that are far away. According to some embodiments, the step of generating the training data can be extended or further defined based on the type of device that collects the media dataset. The device type data can be implemented using adaptive diffusion as described below.

[0048] Once appropriate training data is collected, the AI / ML training system generates a new ANN model and trains it according to the analysis criteria. The AI / ML training system of the platform may store the trained ANN model and other models in the data storage device for reading when required. The step of storing the trained model may further include the step of storing the associated training metadata and the associated analysis criteria (e.g., configuration and parameters), both of which may be included as use case data. According to some embodiments, the use case data may indicate how a particular model can be used and / or what the purpose of such a model can be. For example, the model may be used to implement selective attention on a media dataset or, further, to extract areas therein.

[0049] According to some embodiments, the AI / ML training system searches the data storage device with respect to the trained model using a meta-architecture capable of best implementing or otherwise handling the data indicated by the use case data. Thus, the data storage device may be searched or otherwise filtered based on the use case data of the stored model (e.g., analysis criteria, metadata for training). According to some embodiments, similar use case data across multiple models may indicate the meta-architecture of the models stored therein.

[0050] For example, an AI / ML training system may search its associated data storage device for an ANN model that is trained to detect objects of a particular size. Thus, the model identified by this search may be defined as a particular meta-architecture that represents an architecture capable of identifying objects of interest at a particular size. Similarly, an AI / ML training system may search its associated data storage device for an ANN model that is trained to analyze the relative complexity of the foreground and background of a received media data set as input. Thus, the model identified by this search may be defined as a particular meta-architecture that represents an architecture capable of analyzing the relative complexity of the foreground and background of a media data set.

[0051] According to some embodiments, the meta-architecture may be further identified or otherwise organized as a custom meta-architecture within the data storage device. The custom meta-architecture may be identified by a use case for a lower-level model, such as a model used for selective attention or a model used for object detection. According to some embodiments, the ANN itself, as well as other trained search models, may be used to perform or otherwise extract results from a search of the data storage device associated with the AI / ML training system. Thus, one or more search ANNs may be used to identify meta-architecture candidates that contain models (or otherwise the models themselves) similar to those of the use case identified by the user. For example, the user may provide an analysis criterion or other data indicating a model for determining the complexity of a media data set to the search ANN, and as a result, the search ANN may return a meta-architecture (or otherwise a model therein) indicating such a use case.

[0052] According to some embodiments, the meta-architecture search used for searching the ANN may be similarly trained according to other ANNs provided by the platform. The search ANN may further be trained according to a unique loss function. For example, the search ANN may be trained using, among other techniques, a selective attention metric. Further, the search ANN may be optimized according to various characteristics required by a specific search, such as, among other things, a specific search order, priority, density, and depth of the search space. Similarly, the search ANN may be optimized using Bayesian optimization strategies, according to Gaussian processes, or otherwise using statistical weighting, and determining the correlation of data associated with analysis criteria (e.g., training cycle parameters) with training data and / or use case data.

[0053] The AI / ML training system may use an ANN to find an optimal error threshold for a particular model and use case according to analysis criteria, such as, among other things, provided by a user in the data. For example, a model with a use case for finding a cluster of pixels (or region of interest or "ROI") representing an object of interest based on a three-dimensional map of the input (e.g., x location, y location, and probability that the object of interest exists at that location) may be given a specific error threshold. Thus, the ANN may determine a higher error threshold for models with similar use cases with additional levels of complexity, such as additional input dimensions (e.g., x location, y location, probability that the object of interest exists at that location, and time index of a particular frame). One approach for identifying regions of interest and objects within those regions is described in U.S. Patent Application No. 16 / 953,585, the entire disclosure of which is incorporated herein by reference.

[0054] In certain embodiments of the present invention, a ground truth polygon mask (or “ground truth label”) may be used to define the ROI within an image. In conventional techniques, a binary decision is made based on pixel values such that pixels inside the polygon are considered part of the object while pixels outside the polygon are considered “not an object.” In certain embodiments of the present invention, a “pixel saliency value” can be assigned as a z-value for each x-y pixel location within the ground truth polygon, representing the likelihood that the pixel is part of the object, and a saliency surface map can be generated from the ROI. In some cases, pixels or groups of pixels that meet a certain likelihood threshold can be inferred to be part of the object.

[0055] In some cases, instead of (or in addition to) calculating the saliency value independently or assigning it to each pixel, the curve shape can be applied to the saliency surface map based on the expected object within the ROI, such as the head shape when a human face is expected. The curve shape associated with "head" (e.g., hat) can be used to make an estimation as to whether the object is a head. In some instances, each pixel can be assigned an initial value based on a predetermined distribution regarding the object, and a difference value may be calculated. For example, face recognition can be best predicted using a "hybrid Gaussian" curve, where the initial gradual increase in saliency occurs at the edge of the ROI, and the values across the ROI follow a Gaussian gradient shape such that pixels closer to the center of the ROI have higher saliency values than those along the edge. In some cases, different curve shapes may be used to infer the presence of different objects of interest within the ROI. For example, for smaller persistent objects such as road signs, a Poisson distribution may be used to assign saliency values to pixels, while a different distribution may be used for larger objects where the edge boundary is important, such as a car or other vehicle. The "fit" between a particular shape (or series of shapes) and the object of interest can then be used to further train the object ANN model for subsequent object detection.

[0056] These gradient values are applied to various images and used as input into the training step based on the fitness and accuracy to further refine each model for specific object use cases, devices, or combinations thereof.

[0057] According to some embodiments, once the model is identified by the AI / ML training system, it may be further trained or otherwise optimized using transfer learning (e.g., adaptive divergence), as described below for example.

[0058] According to some embodiments, the AI / ML training system may further traverse the remote hardware of the customer environment and, in some cases, use and merge two or more models to determine a model collaboration structure and facilitate the intelligent distribution of data. For example, the model used for selective attention may be structured or otherwise arranged to transfer the feature map to a second model, such as an object detection model, rather than acting on the image data it first receives. This deployment option may be useful, for instance, in an instance where an initial selective attention model and a second object detection model are combined into a single wrapper function and provided via control flow software to an endpoint device. In such an example, a "switchable pipeline" implementation may be used, where two (or in some cases, more than two) models can be executed on the same input data either in parallel or sequentially, as directed by a user or a pre-configured switch. Thus, a device that may only support the execution of a single model due to processor and / or power constraints can perform two distinct but "merged" ANN model estimations (e.g., finding the area of interest in an image and then the objects within the area of interest) using two different models.

[0059] The model collaboration structure, as determined by the AI / ML training system, may be stored by a data storage device associated with the business logic component of the platform, or otherwise indicated.

[0060] According to some embodiments, the platform may be constructed from an SDK and other software from a single ANN model as provided by an AI / ML training system. Thus, a compiler as implemented by the platform can compile models for multiple different hardware architecture targets for execution on specific hardware located within the customer environment. In conventional implementations, ANN models are trained using hardware-independent parameters. This approach simplifies training and deployment but suffers from accuracy and performance. By doing so, when the model is compiled, minor changes to the processing can be introduced, which can result in near-optimal processing on a certain hardware device. To address this problem, in some embodiments of the present invention, the training process is executed on a processor (or emulator) specific to the edge device hardware (e.g., a specific model of camera). Analyzing the results of the training steps that occur on the specific hardware enables the model to be trained for that specific device, resulting in a hardware-specific variant of the model optimized for the processor used within that device. In some cases, an "emulator library" is provided for each specific device to process the training data and model.

[0061] According to some embodiments, the SDK and other software are distributed using a unique DRM system. In one example, the unique DRM system provides a heartbeat-like system, and data can be transmitted to a central licensing server based on a predetermined time period. The data transmitted during each beat of the heartbeat-like system may include, but is not limited to, detection event data such as scene indications, location (e.g., GPS coordinates), width, height, and time index, estimated speed, frequency of SDK usage, and other data associated therewith, media data sets and data and metadata of the image itself. The predetermined time period may be unique to each user environment or each device. In some embodiments, a risk engine may be incorporated into the unique DRM system to determine whether a license associated with a customer environment utilizes a longer or shorter predetermined time period. The risk engine may also determine whether a license should be denied based on one or more license restrictions, usage limits, time periods, etc. Further, data received by the DRM system from a device or other endpoint that is suspected or otherwise considered suspicious may be brought to the user's attention.

[0062] According to some embodiments, the unique DRM system may further track how software and other data are used by the hardware of the customer environment. For example, the DRM system may track, among other data, the use of specific models, detections associated with each model, and use case detections. Thus, the DRM system can determine a price for the user based on the software and data usage amounts tracked on each device or hardware in the user environment.

[0063] According to some embodiments, the data provided within the heartbeat-like system may be used to identify a device that is malfunctioning or a device that has been tampered with or disrupted through a heuristic or risk engine type system.

[0064] The platform further provides various artificial intelligence (AI) detection mechanisms to the Visual Intelligence SDK. According to some embodiments, the Visual Intelligence SDK may include features such as, among others, estimation, image processing, unique DRM, quality control sampling, and wireless updates. The Visual Intelligence SDK may implement a dynamic post-processing analysis engine and detect shared terms across one or more sequential images from one or more devices across the user environment.

[0065] Similarly, the shared term detection of the Visual Intelligence SDK may be implemented in the same way across one or more user environments to detect shared terms across one or more user environments. In such cases, the process, prior to edge deployment, may increase the number of frames and voting strategies to be classified across multiple frames of an image so that discrepancies are reduced and thus the accuracy of the model can be improved. More specifically, the user can define a temporal window (e.g., 10 frames) of frames from a video file to perform multiple estimations on some or all of the frames within that window. The comparison is performed to measure the difference between an image known to contain the object (e.g., the face of a particular person) and the images within each frame. If a certain percentage of the frame embeddings (e.g., 50%, which can be user-defined) fall below a specified distance threshold (which can also be user-defined), the object is considered to be the same as the object captured from multiple frames.

[0066] The Visual Intelligence SDK may also implement privacy features such as those provided by the platform's Privacy ANN or other artificial intelligence models. For example, a model trained for selective attention may be implemented to filter out information that requires attention in handling (e.g., face, PPI, nude, or other data that requires attention in handling) for use in quality assurance (QA) sampling by the platform. Similarly, a model trained for selective attention may be distributed within edge devices or other hardware within the user environment to filter out information that requires attention in handling before further analysis or transmission to other devices. The model may perform such privacy filtering by encrypting, revising, obfuscating, or compressing the field of the media dataset, or alternatively, by using cropping features to remove data that requires attention in handling. The model may also extract specific areas (e.g., objects of interest) from the media dataset for privacy purposes (removing the rest of the dataset). For example, after identifying an object of interest within the media dataset, one or more models may extract only the object of interest, remove the rest of the environment within the field of view, and maintain other privacy within the rest of the environment. Thus, the smaller image extracted by the model further includes data annotations therein, identifies the location of the extracted image within the original field of view, and facilitates the construction of the original media dataset from which unimportant data has been removed.

[0067] In some embodiments, privacy filtering may encrypt a view or a portion of a view using techniques that require multiple factors for decoding. In some embodiments, one such factor is a unique token that changes over time. In such embodiments, a media image or video may be decoded by using only user authentication factors in conjunction with a specific token corresponding to a certain time period on a specific device or group of devices. In some embodiments, the decoding token is digitally stored such that the reading of the token is recorded, for example, for auditing or control purposes.

[0068] Unique data structure (for transmission of video data with variable resolution) Similar to the privacy features implemented by the Visual Intelligence SDK detailed above, the Visual Intelligence SDK (or other software provided by the platform) may also extract specific areas (e.g., the object of interest) from the media data set (removing the rest of the data set) in order to generate a smaller media data set of images and reduce the file size for transmission. Thus, the smaller images extracted by the model may further include data annotations therein, identify the location of the extracted images within the original field of view, and facilitate the construction of the original media data set from which uninteresting data has been removed. By removing uninteresting data from the media data set prior to transmission, a smaller media data set can be transmitted via the network at a significantly reduced file size. Referring to FIG. 2, a data file architecture (205) representing video data using the H.264 protocol contains a significant amount of data not required for image detection and extraction by the techniques described herein. Instead, a slice (or slices) of data (210) is selected to contain content that encapsulates one or more regions of interest from the video segment. The content of the "slice" 210 may include various parameters 215 related to the region of interest within the image. The parameters may include, for example, the start and end of the time index, the upper, lower, right, and left coordinates, as well as the extracted image itself, or, in some cases, a downsampled version of the image.

[0069] The reduced file size may further be accomplished by transmitting, at the edge device of the user environment, an area of the media data set that contains the object of interest at high resolution while transmitting the remainder of the field of view at low resolution. The high and low resolution areas of the digital media data set may be transmitted as a reduced media data set with a file size significantly smaller than the original media data set. Similar to the reconstruction described above, the reduced media data set may further include data annotations therein, identify the locations of the high and low resolution images within the original field of view, and facilitate the construction of the original media data set with areas containing only the objects of interest at high resolution. The video file may repeat the file reduction and construction process for each frame of the video file.

[0070] The AI / ML training system of the platform as described herein may further provide the ability for a pre-trained model to be further trained using data associated with real-time data collected by edge devices and other hardware of the user's environment. As described above, once a model is identified by the AI / ML training system, it may be further trained or otherwise optimized using a training method called adaptive divergence (e.g., a persistent transfer learning method). Alternatively, adaptive divergence may be implemented on an ANN model that is already distributed across the hardware of the user's environment, e.g., by means of federated learning techniques.

[0071] According to some embodiments, and referring to FIG. 3, the user's environment may contain an ANN model that is supplied based on, or alternatively from, the data storage device of the AI / ML training system 120 of the platform 100. Quality assurance (QA) sampling data (“training data”) is collected from the model and transmitted to the auditing module (step 305) and may be used as an initial training set for the initial ANN / AI model (step 310). The model is then deployed and results can be collected from its use within the field (step 315). The QA sampling data may be audited using an automated auditing process and / or human auditing (e.g., audited by a user) (step 320). The automated auditing process may be performed by the ANN model to identify relevant (e.g., correct) data from the QA sampling data. Alternatively, a human (e.g., a user) may audit the QA sampling data for accuracy and manually mark the results as either correct or incorrect. Similarly, a human may audit the QA data in conjunction with the ANN model to facilitate human auditing of the QA sampling data. The audited QA sampling data (e.g., data marked as correct) is stored as training data or alternatively incorporated into the training data stored in the data storage device for applying updates to the ANN model during retraining. Once the training data is updated, an updated ANN model may be generated or alternatively trained therefrom (step 325). The platform may distribute the updated ANN model to the hardware of the user's environment to provide a better trained model for the user's specific use case.

[0072] According to some embodiments, and referring to FIG. 4, the adaptive divergence procedure as described above may be provided for different devices of the hardware of the user environment or, alternatively, implemented differently. For example, QA sample data may be vetted by an ANN model and, after vetting, provided to a data storage device that is associated with the individual training data of a particular device. Similar to the training described above, the updated training data may be vetted by an ANN model, a human, or a combination thereof. The vetted QA sample data stored within the data storage device may be correlated with feedback data or other media information associated with a media data set, such as, among other things, brightness, background complexity, size or geometric shape of the object of interest, and data associated with the device (e.g., statistical device information). The ANN model may select the vetted QA sample data and retrain or further update the ANN model for one or more target devices. Thus, each particular device may receive updated or further updated model training using only the vetted QA sample data received from that particular device. Alternatively, the QA sample data may be provided to a data storage device that is associated with the training data for a particular group of devices, and the group of devices may be enabled to receive a trained and further updated model using only the vetted QA sample data received from that particular group of devices.

[0073] More specifically, QA sample data is collected and / or received from various sources representing specific use cases (step 405), and an initial AI model is trained using the collected data (step 410). As with previous use cases, the AI model is deployed into the field corresponding to the use case, and results are collected (step 415). In this embodiment, the results can be collected and stored separately, or alternatively, can be identified as coming from different devices or deployments (step 420), while in other instances, the data can be grouped into a single dataset. The results are then vetted for accuracy using either an automated process, human scrutiny, or a combination of both (step 425). Once deemed accurate and sufficient, the images and associated results are corrected and used to create an updated training dataset (step 430), and in one instance, the dataset can be segmented such that specific datasets are assigned to specific devices or groups of devices (step 440). The grouping can be based on several commonalities such as the camera manufacturer, model, and / or model, the environment in which the device is used (e.g., training sets regarding images installed outdoors vs. indoors, night-time images vs. daytime images, etc.), and / or functional use case commonalities such as face recognition vs. character recognition. Once updated, the training datasets are created for specific devices, and they are then used to train an updated AI model for the specific device (step 450). The process then iterates over time as new AI models are deployed and used in the field, new data is collected, and the process is repeated, resulting in an AI model that is continuously improved for a specific device and / or use case.

[0074] In addition to the training described above, the platform may further provide a retraining maintenance module that monitors models for all devices in the user environment and automatically retrains each of them. The retraining maintenance module may incorporate an ANN or other model and determine the monitoring criteria associated therewith. For example, the retraining maintenance module may use Bayesian optimization to determine how often the retraining maintenance module reads QA sample data for quality assurance and quality control purposes. Further, the retraining maintenance module may further determine, using the ANN model or other model, how the devices in the user's environment should be grouped or otherwise organized for training. For example, the retraining maintenance module may optimize the size of each device group based on its use case or other data associated with the device. Devices may share vetted QA sample data (e.g., auxiliary training data) for each device group to most effectively train each ANN model therein.

[0075] Implementations of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in combinations of one or more of them, including the structures disclosed in this specification and their structural equivalents. Implementations of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiver apparatus suitable for execution by a data processing apparatus.

[0076] A computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or can be included therein. Further, a computer storage medium is not a propagated signal, but a computer storage medium can be a source or destination of computer program instructions encoded within an artificially generated and propagated signal. A computer storage medium can also be or can be included in one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0077] The operations described herein can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0078] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for data processing, and by way of example, includes programmable processors, computers, systems on a chip, or a plurality or combination of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program, e.g., processor firmware, protocol stack, database management system, operating system, cross-platform runtime environment, virtual machine, or a combination of one or more of them. The apparatus and the execution environment can implement various different computing model infrastructures, e.g., web services, distributed computing, and grid computing infrastructures.

[0079] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiler-type or interpreter-type languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (such as one or more scripts stored in a markup language resource), in a single file dedicated to the program, or in multiple cooperating files (such as files that store one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers, located in one place or distributed across several locations and interconnected by a communication network.

[0080] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs, acting on input data and performing actions by generating output. The processes and logical flows can also be performed by special purpose logic circuits, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as such.

[0081] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. Essential elements of a computer include a processor for performing actions in accordance with instructions, and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as, magnetic, magneto-optical disks, or optical disks, or will be operatively coupled to receive data therefrom, transfer data thereto, or both. However, a computer need not have such devices. Further, a computer may be embedded within another device, such as, by way of example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable memory device (e.g., a universal serial bus (USB) flash drive). Devices suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, such as, EPROM, EEPROM, and flash memory devices, magnetic disks, such as, internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0082] The terms and phrases used in this specification are for the purpose of explanation and should not be regarded as limiting. Indefinite articles "a" and "an" as used in this specification and the claims should be understood to mean "at least one" unless clearly indicated otherwise. The phrase "and / or" as used in this specification and the claims should be understood to mean "either or both" of the elements so combined, i.e., elements that in some cases coexist conjunctively and in other cases exist disjunctively. Multiple elements listed using "and / or" should be construed in the same manner, i.e., as "one or more than one" of the elements so combined. Other elements not specifically identified by the "and / or" clause may optionally exist, regardless of whether they are related to those specifically identified elements. Thus, by way of non-limiting example, when used in conjunction with non-limiting language such as "comprising", a reference to "A and / or B" can, in one embodiment, refer to only A (optionally including elements other than B), in another embodiment, refer to only B (optionally including elements other than A), and in yet another embodiment, refer to both A and B (optionally including other elements), etc.

[0083] As used in this specification and the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be construed as inclusive, i.e., as including some elements or at least one of the list of elements (including those greater than one) and optionally additional unlisted items. Only terms such as "only one of", "exactly one of", or when used in the claims, terms such as "consisting of" where the contrary is clearly indicated, will refer to the inclusion of exactly one of some elements or the list of elements. Generally, the term "or" as used will be construed to indicate exclusive alternatives (i.e., "not both, but one or the other") only when preceded by exclusive terms such as "either", "one of", "only one of", or "exactly one of". When used in the claims, "consisting essentially of" shall have its ordinary meaning as used in the field of patent law.

[0084] As used in this specification and the claims, the phrase "at least one" in reference to a list of one or more elements should be understood to mean at least one element selected from one or more of the elements in the list of elements, but does not necessarily include at least one of every element specifically recited in the list of elements, nor does it exclude any combinations of elements in the list of elements. This definition also allows for the possibility that there may optionally be elements other than those specifically identified in the list of elements that the phrase "at least one" refers to, whether or not they are related to those specifically identified elements. Thus, by way of non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B" or equivalently, "at least one of A and / or B") can, in one embodiment, refer to at least one A (optionally including more than one A and optionally including elements other than B) in the absence of any B, in another embodiment, refer to at least one B (optionally including more than one B and optionally including elements other than A) in the absence of any A, and in yet another embodiment, can refer to at least one A (optionally including more than one A) and at least one B (optionally including more than one B) and optionally other elements, etc.

[0085] The use of "including", "comprising", "having", "containing", "involving", and variations thereof means including the items listed thereafter and additional items.

[0086] The use of ordinal terms such as "first", "second", "third", etc. in a claim for purposes of amending claim elements does not, by itself, imply any priority, precedence, or order of one claim element with respect to another, or the temporal order in which acts of a method are performed. Ordinal terms are used merely to distinguish one claim element having a certain name from another element having the same name (in the absence of the use of ordinal terms) and are used as labels for distinguishing claim elements.

[0087] Features described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, for the sake of brevity, various features described in the context of a single embodiment may also be provided separately or in any suitable sub-combination. The Applicant hereby notifies that new claims may be formulated against such features and / or combinations of such features during the patent examination of the present application or any further application derived therefrom. The features of the devices and systems described are incorporated into / used in the corresponding methods and vice versa.

Claims

1. A method for generating a plurality of artificial neural network (ANN) models specific to a device for distribution across a plurality of user devices, the method comprising: a processor receiving a plurality of sample data sets from the plurality of user devices in the user environment, the plurality of sample data sets comprising media data and a plurality of predictions by a plurality of ANN models specific to the devices executed on the plurality of user devices, each of the plurality of ANN models specific to the devices being specifically developed for a corresponding one of the plurality of user devices, and the plurality of predictions generated by the plurality of ANN models specific to the devices comprising a per-image quantitative image saliency metric indicating the likelihood that a media file contains a target object; the processor writing the plurality of sample data sets to a training data storage device; the processor identifying, within a data storage device, a use case data set, the use case data set comprising at least a plurality of training data parameters, the plurality of training data parameters comprising device parameters, the device parameters including one or more of available memory, processing speed, image resolution, and capture frame rate of an edge device; the processor identifying, within the training data storage device, training data from the plurality of sample data sets that satisfies the plurality of training data parameters provided within the use case data set; the processor identifying, within the data storage device, a plurality of stored instances of the plurality of ANN models specific to the device; the processor generating, based on the training data, a plurality of updated device-specific ANN models from each of the plurality of stored instances of the plurality of device-specific ANN models; A method, including the above steps.

2. The method according to claim 1, wherein the plurality of user devices comprises a plurality of image capture devices.

3. The media data includes image data, and applying a plurality of ANN models specific to the device to the image data promotes identifying an object of interest in the image data. The method according to claim 1 or claim 2.

4. The plurality of training data parameters include media data parameters. The method according to claim 1 or claim 2 or claim 3.

5. The media data parameters include one or more of a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and a gamma value. The method according to claim 4.

6. The use case dataset is specific to a particular use case. The method according to claim 1 or any one of claims 2 to 5.

7. The use case includes an environmental aspect and a functional aspect. The method according to claim 6.

8. The functional aspect of the use case includes face recognition. The method according to claim 7.

9. The environmental aspect of the use case includes one of the installation of the device in an outdoor environment, the installation of the device in an indoor environment, the installation of the device in a well-lit environment, or the installation of the device in a poorly-lit environment. The method according to claim 7 or claim 8.

10. The method further includes determining a minimum number of images required to achieve a threshold model accuracy, at least partially based on the quantitative image saliency metric for each of the images. The method according to claim 1 or any one of claims 2 to 9.

11. The method further includes maintaining a library of a plurality of device-specific parameters and training data, and generating the updated plurality of ANN models specific to the device is further based on the plurality of device-specific parameters and training data. The method according to claim 1 or any one of claims 2 to 10.

12. The method further includes the processor distributing the updated plurality of ANN models specific to the device to at least a subset of the plurality of user devices associated therewith. The method according to claim 1 or any one of claims 2 to 11.

13. A system for generating a plurality of artificial neural network (ANN) models specific to devices for distribution across a plurality of user devices, the system comprising: one or more processors; a memory coupled to the one or more processors; wherein the one or more processors execute a plurality of modules stored in the memory; the plurality of modules include a user interface that receives instructions from a user, the instructions identifying one or more sample data sets from the plurality of user devices of the user environment, the one or more sample data sets comprising media data and a plurality of predictions by a plurality of ANN models specific to the devices executed on the plurality of user devices, each of the plurality of ANN models specific to the devices being specifically developed for a corresponding one of the plurality of user devices, and the plurality of predictions generated by the plurality of ANN models specific to the devices comprising a per-image quantitative image saliency metric indicating the likelihood that a media file contains a target object, the user interface; a data storage device comprising the one or more sample data sets; a business logic module that, when executed, (i) identifies a use case data set stored in the data storage device, the use case data set comprising at least a plurality of training data parameters, the plurality of training data parameters comprising device parameters, the device parameters including one or more of available memory, processing speed, image resolution, capture frame rate of an edge device, (ii) identifies training data from the one or more sample data sets that satisfies the plurality of training data parameters provided in the use case data set, and (iii) identifies the plurality of ANN models specific to the devices stored in the data storage device, the business logic module; An artificial intelligence machine learning module, which, when executed, generates a plurality of ANN models specific to the device, updated from each of a plurality of stored instances of ANN models specific to the device, based on the training data. The artificial intelligence machine learning module A system comprising the same.

14. The system according to claim 13, wherein the plurality of user devices include a plurality of image capture devices.

15. The system according to claim 14, wherein the media data includes image data, and applying the plurality of ANN models specific to the device to the image data facilitates identifying the object of interest in the image data.

16. The system according to claim 13, 14, or 15, wherein the plurality of training data parameters include media data parameters.

17. The system according to claim 16, wherein the media data parameters include one or more of a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and a gamma value.

18. The system according to claim 13 or any one of claims 14 to 17, wherein the use case dataset is specific to a particular use case.

19. The system according to claim 18, wherein the use case includes an environmental aspect and a functional aspect.

20. The system according to claim 19, wherein the functional aspect of the use case includes face recognition.

21. The system according to claim 19 or 20, wherein the environmental aspect of the use case includes one of installation of the device in an outdoor environment, installation of the device in an indoor environment, installation of the device in a well-lit environment, or installation of the device in a poorly-lit environment.

22. The plurality of predictions generated by the plurality of ANN models specific to the device include a per-image quantitative image saliency metric indicating the likelihood that the media file contains the object of interest. The system according to claim 13 or any one of claims 14 to 21.

23. The system according to claim 22, wherein the artificial intelligence machine learning module further determines a minimum number of images necessary to achieve a threshold model accuracy, based at least in part on the quantitative image saliency metric for each of the images. **Claim 24** The system according to any one of claims 13 or 14 to 23, further comprising a library of a plurality of device-specific parameters and training data, wherein the artificial intelligence machine learning module generates a plurality of ANN models specific to the updated device based on the plurality of device-specific parameters and the training data. **Claim 25** The system according to any one of claims 13 or 14 to 24, further comprising a deployment module for distributing the plurality of ANN models specific to the updated device to at least a subset of the plurality of user devices associated therewith.

Citation Information

Patent Citations

  • Imaging device, image processing method, and program thereof

    JP2010147696A

  • Method and device for handling multiple video streams using metadata

    JP2013176101A

  • Object detection device and object detection method and program

    JP2019008460A

  • Server for learning, image collection assisting system for insufficient learning, and image estimation program for insufficient learning

    JP2019192082A

  • Server device, trained model providing program, trained model providing method, and trained model providing system

    WO2018173121A1