Systems and methods for distributed data analysis

By distributing device-specific ANNs across user devices, the system addresses computational intensity in data analytics, improving data processing efficiency and accuracy for diverse datasets.

JP7781125B2Active Publication Date: 2025-12-05ゼイリエント
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023178050
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-08
Filing Date
2023-10-16
Publication Date
2025-12-05
Estimated Expiration
2041-05-10

AI Technical Summary

Technical Problem

Existing data analytics platforms face challenges in efficiently processing large and diverse datasets due to computational intensity, requiring improved methods and systems for distributing data analysis across multiple devices using neural networks.

Method used

The system facilitates complex data analysis by distributing device-specific artificial neural networks (ANNs) across user devices like smartphones and IoT devices, utilizing a unified platform for data processing, user interface, and model training, with features like image saliency metrics and curve shape fitting to identify objects of interest.

Benefits of technology

This approach enhances data availability, training effectiveness, and execution efficiency, enabling broader adoption and more accurate dataset utilization for user devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007781125000001
    Figure 0007781125000001
  • Figure 0007781125000002
    Figure 0007781125000002
  • Figure 0007781125000003
    Figure 0007781125000003
Patent Text Reader

Abstract

To provide a system and a method for generating a device-specific Artificial Neural Network (ANN) for distribution across user devices (105).SOLUTION: Sample datasets (140) are collected from devices in a particular environment or use case and include predictions by device-specific ANN models to be executed on the user devices. The received datasets are used with existing datasets and stored ANN models to generate updated device-specific ANN models from each of the stored instances of the device ANN models based on the training data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 021,735, entitled "Systems and Methods for Distributed Data Analytics," filed May 8, 2020, the entire disclosure of which is incorporated herein by reference.

[0002] The following disclosure is directed to methods and systems in data analysis, and more particularly to a distribution of data analysis frameworks and associated data tools. [Background technology]

[0003] With the development of intelligent systems, the amount of data that is retrieved, transmitted and further processed is continually increasing.

[0004] Complex data analysis may implement machine learning mechanisms and use large training datasets to train neural networks. These neural networks may then be used to process input data in domains similar to the large training dataset. Applying such complex neural network logic to broader and larger datasets of input data may prove difficult and computationally intensive. Therefore, as disclosed herein, methods and systems for providing access to data processing tools and distributable data analysis platforms provide such systems with the benefits of higher platform adoption and distribution rates, greater data availability, improved training effectiveness, and improved execution efficiency.

[0005] Thus, improved methods and systems for data processing using neural networks can significantly benefit from improved execution efficiency. Summary of the Invention [Means for solving the problem]

[0006] Current data analytics platforms use a variety of external tools to accomplish specific tasks. This disclosure describes techniques and related systems that facilitate complex data analysis using tools that can be accessed remotely or otherwise distributed across multiple devices using a defined analytics framework in a uniform platform.

[0007] Disclosed herein are exemplary embodiments of methods and support systems that facilitate data analysis using a wide range of data stores, neural networks, and other data science tools. Providing access to such powerful tools supports a broader adoption base and therefore results in larger and more accurate datasets for training and development purposes. Ease of user interaction depends primarily on the user interface provided by the platform and the methods used to provide access to the powerful data analysis tools, model training mechanisms, and neural networks implemented therein. The various systems and methods provided by the present invention actively employ multiple neural networks that work in conjunction to process input data and identify zones of interest and / or objects of interest within the zones, while providing users with an easy-to-read and readable interface to review the input and training data, improve the accuracy of the models used by the networks, and visualize the results.

[0008] Objects can be inanimate objects and generally identified (e.g., "cars" or "pedestrians"), or concrete objects and specifically identified, for example, based on a combination of facial recognition, character recognition, or similar techniques.

[0009] Thus, in a first aspect, the present invention provides a method for generating device-specific artificial neural network (ANN) models for distribution across user devices, such as smartphones, cameras, and other Internet of Things (IoT) devices. In various embodiments, the method includes receiving, by a processor, a sample dataset from a user device in a user environment, the sample dataset comprising media data and predictions by a device-specific ANN model executing on the user device, and writing, by the processor, the sample dataset to a training data store. The method also includes identifying, by the processor, in the data store: (i) a use case dataset, the use case dataset comprising at least training data parameters; (ii) training data from the sample dataset that satisfies the training data parameters provided in the use case dataset; and (iii) stored instances of the device-specific ANN model. The processor then generates updated device-specific ANN models from each stored instance of the device ANN model based, at least in part, on the identified training data. In some cases, a library of device-specific parameters and training data is maintained, and the generation of the updated device-specific ANN model is further based on the device-specific parameters and the training data.

[0010] In some embodiments, the media data comprises image data, and application of the ANN model to the image data facilitates identification of objects of interest within the image data. The training data parameters may include media data parameters such as a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and / or a gamma value, and / or device parameters such as available memory, processing speed, image resolution, and / or capture frame rate.

[0011] In some cases, the use case dataset is specific to a particular use case and may, in some instances, include environmental aspects (such as device placement in an outdoor environment, device placement in an indoor environment, device placement in a well-lit environment, or device placement in a poorly lit environment) and functional aspects (e.g., facial recognition, character recognition, document authentication, etc.). In some embodiments, the predictions generated by the device-specific ANN model include a per-image quantitative image saliency metric indicating the likelihood that a media file contains an object of interest, and in some cases determine a minimum number of images required to achieve a threshold model accuracy based, at least in part, on the per-image quantitative image saliency metric.

[0012] In some cases, the method further includes distributing the updated device-specific ANN model to at least a subset of the user devices associated therewith.

[0013] In another aspect, the present invention provides a system for generating device-specific artificial neural network (ANN) models for distribution across user devices, such as smartphones, cameras, and other Internet of Things (IoT) devices. The system includes one or more processors and a memory coupled to the processor, the processor executing a plurality of modules stored in the memory. The modules include a user interface that receives instructions from a user to identify one or more sample datasets from a user device in the user's environment, the sample dataset comprising media data and predictions by a device-specific ANN model executing on the user device; a data store that comprises the sample dataset; a business logic module that, when executed, (i) identifies a use case dataset to be stored in the data store, the use case dataset comprising at least training data parameters; (ii) identifies training data from the sample dataset that satisfies the training data parameters provided in the use case dataset; and (iii) identifies a device-specific ANN model to be stored in the data store; and an artificial intelligence machine learning module that, when executed, generates an updated device-specific ANN model from each stored instance of the device ANN model based on the training data.

[0014] In some embodiments, the media data comprises image data, and application of the ANN model to the image data facilitates identification of objects of interest within the image data. The training data parameters may include media data parameters such as a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and / or a gamma value, and / or device parameters such as available memory, processing speed, image resolution, and / or capture frame rate.

[0015] In some cases, the use case dataset is specific to a particular use case and, in some instances, may include environmental aspects (such as device placement in an outdoor environment, device placement in an indoor environment, device placement in a well-lit environment, or device placement in a poorly lit environment) and functional aspects (e.g., facial recognition, character recognition, document authentication, etc.). In some embodiments, the predictions generated by the device-specific ANN model include a per-image quantitative image saliency metric indicating the likelihood that a media file contains an object of interest, and in some cases determine a minimum number of images required to achieve a threshold model accuracy based, at least in part, on the per-image quantitative image saliency metric.

[0016] In some cases, the distribution module distributes the updated device-specific ANN model to at least a subset of the user devices associated therewith.

[0017] In another aspect, the present invention provides a method for optimizing execution of device-specific trained artificial neural network (ANN) models on edge devices (such as smartphones, cameras, and other Internet of Things (IoT) devices), comprising: receiving, by a processor, a first trained ANN model and a second ANN model, wherein the first ANN model and the second ANN model each perform different inferences on input data, and an output of the first ANN model serves as an input to the second ANN model; and merging, in accordance with control flow instructions, the first ANN model, the second ANN model, and the control flow execution instructions into a combined software package for execution thereon and for deployment to the edge device.

[0018] In an embodiment, the first trained ANN model and the second trained ANN model each comprise distinct analytical criteria and use case data, and the processor selects the first and second ANN models based at least in part on the analytical criteria therein. A parent ANN may be generated as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and the meta-architecture can then be delivered to an edge device such that it executes as a single ANN model. In an embodiment, the edge device is a camera, and execution of the first ANN model and the second ANN model on the camera can identify an object of interest in an image file captured on the camera.

[0019] In another aspect, the present invention provides a system for optimizing the execution of device-specific trained artificial neural network (ANN) models on edge devices (such as smartphones, cameras, and other Internet of Things (IoT) devices). The system includes one or more processors and a memory coupled to the one or more processors, wherein the one or more processors execute computer-executable instructions stored in the memory. When executed, the instructions identify, in a data storage device, a first trained ANN model and a second ANN model, each of the first ANN model and the second ANN model performing a different estimation on input data, an output of the first ANN model serving as an input to the second ANN model, merging the first ANN model, the second ANN model, and control flow execution instructions into a combined software package, and deploying, using a distribution module, the combined software package to the edge device for execution thereon in accordance with the control flow instructions.

[0020] In an embodiment, the first trained ANN model and the second trained ANN model each comprise distinct analytical criteria and use case data, and the processor selects the first and second ANN models based at least in part on the analytical criteria therein. A parent ANN may be generated as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and the meta-architecture can then be delivered to an edge device such that it executes as a single ANN model. In an embodiment, the edge device is a camera, and execution of the first ANN model and the second ANN model on the camera can identify an object of interest in an image file captured on the camera.

[0021] In another aspect, the present invention provides a method for identifying an object of interest in image files. The method includes receiving one or more image files, each image file potentially containing an object of interest, and applying a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values ​​indicating the likelihood that a particular pixel is part of the object of interest. Based on the ground truth labels, a three-dimensional saliency surface map having an x-axis, a y-axis, and a z-axis is generated, where the x-axis and y-axis values ​​define the location of the pixel in the image and the z-axis value is the pixel-specific saliency value. A curve shape is selected from a library of curve shapes, the curve shape is applied to the saliency surface map, a fit between the curve shape and the three-dimensional surface is determined, and based on the fit, it is determined whether the image file contains the object of interest.

[0022] In some embodiments, the curve shape is selected based on the object of interest and may be based, at least in part, on one or more statistical distributions, such as a Gaussian distribution, a Poisson distribution, or a hybrid distribution. In some cases, the image file is added to a library of image files for use in training an artificial neural network (ANN), and the ANN may be trained to identify the object of interest in subsequent media files and / or segment objects in subsequent media files.

[0023] In another aspect, the present invention provides a system for identifying an object of interest in an image file, the system including one or more processors and a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory. When executed, the system receives one or more image files, each image file potentially containing an object of interest, and applies a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values ​​indicating the likelihood that a particular pixel is part of the object of interest. Based on the ground truth labels, a three-dimensional surface having an x-axis, a y-axis, and a z-axis is generated, where the x-axis and y-axis values ​​define the location of the pixel in the image and the z-axis value is the pixel-specific saliency value. A curve shape is selected from a library of curve shapes, the curve shape is applied to the ground truth label, a fit between the curve shape and the three-dimensional surface is determined, and based on the fit, the system determines whether the image file contains the object of interest.

[0024] In some embodiments, the curve shape is selected based on the object of interest and may be based, at least in part, on one or more statistical distributions, such as a Gaussian distribution, a Poisson distribution, or a hybrid distribution. In some cases, the image file is added to a library of image files for use in training an artificial neural network (ANN), and the ANN may be trained to identify the object of interest in subsequent media files and / or segment objects in subsequent media files.

[0025] In yet another aspect, the present invention provides a method for storing image data for transmission of video data, comprising the steps of receiving video data in a standard video data format (such as H.264) at an edge device, and extracting image slices from the video data, the image slices comprising an image, a start index time and an end index time indicating the temporal location of the image slice within the video data, and region of interest parameters describing the two-dimensional coordinates of a region of interest within the image.

[0026] In some embodiments, receiving the video data and extracting the image slices is performed on an edge device. The image slices may then be analyzed using one or more artificial neural networks on the edge device to determine a region of interest and whether the region of interest includes an object of interest. In some cases, the image slice is identified as high-resolution if the image slice includes the object of interest, and as low-resolution otherwise. The method may further include transmitting the high-resolution image slices to an artificial intelligence machine learning module for inclusion in a training dataset specific to the edge device on which the image was captured within the artificial neural network.

[0027] In another aspect, the present invention provides a system for storing image data for transmission of video data, the system including one or more processors and a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory. When the instructions are executed, the system receives video data at an edge device in one of a plurality of standard video data formats (e.g., H.264) and extracts image slices from the video data, the image slices comprising an image, start index times and end index times indicating a temporal location of the image slice within the video data, and region-of-interest parameters describing two-dimensional coordinates of a region of interest within the image.

[0028] In some embodiments, receiving the video data and extracting the image slices is performed on an edge device. The image slices may then be analyzed using one or more artificial neural networks on the edge device to determine a region of interest and whether the region of interest includes an object of interest. In some cases, the image slice is identified as high-resolution if the image slice includes the object of interest, and as low-resolution otherwise. The method may further include transmitting the high-resolution image slices to an artificial intelligence machine learning module for inclusion in a training dataset specific to the edge device on which the image was captured within the artificial neural network.

[0029] Features described in the context of separate aspects and / or embodiments of the invention may be used together and / or interchangeable where possible. Similarly, where features are described, for brevity, in the context of a single embodiment, those features may also be provided separately or in any suitable subcombination. Features described in the context of a system may have corresponding features definable and / or combinable with respect to a method, or vice versa, and these embodiments are specifically contemplated. The present invention provides, for example, the following. (Item 1) 1. A method for generating device-specific artificial neural network (ANN) models for distribution across user devices, the method comprising: receiving, by a processor, a sample data set from the user device in a user environment, the sample data set comprising media data and predictions from a device-specific ANN model executing on the user device; writing, by the processor, the sample data set to a training data store; identifying, by the processor, a use case data set in a data store, the use case data set comprising at least training data parameters; identifying, by the processor, in the training data store, training data from the sample data set that satisfies training data parameters provided in the use case data set; identifying, by the processor, in the data storage device, a stored instance of the device-specific ANN model; generating, by the processor, an updated device-specific ANN model from each stored instance of the device ANN model based on the training data; A method comprising: (Item 2) Item 10. The method of item 1, wherein the user device comprises a plurality of image capture devices. (Item 3) 3. The method of claim 1 or 2, wherein the media data comprises image data, and application of the ANN model to the image data facilitates identification of an object of interest within the image data. (Item 4) 4. The method of claim 1, wherein the training data parameters include media data parameters and device parameters. (Item 5) Item 5. The method of item 4, wherein the media data parameters include one or more of a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and a gamma value. (Item 6) 6. The method of claim 4 or 5, wherein the device parameters include one or more of available memory, processing speed, image resolution, and capture frame rate. (Item 7) 7. The method of claim 1 or any of claims 2-6, wherein the use case dataset is specific to a particular use case. (Item 8) 8. The method of claim 7, wherein the use case comprises an environmental aspect and a functional aspect. (Item 9) Item 9. The method of item 8, wherein the functional aspects of the use case comprise facial recognition. (Item 10) 10. The method of claim 8 or 9, wherein the environmental aspect of the use case comprises one of installing the device in an outdoor environment, installing the device in an indoor environment, installing the device in a well-lit environment, or installing the device in a poorly lit environment. (Item 11) 11. The method of any of items 1 or 2-10, wherein the predictions generated by the device-specific ANN model comprise a quantitative image saliency metric per image indicating the likelihood that the media file contains an object of interest. (Item 12) 12. The method of claim 11, further comprising determining a minimum number of images required to achieve a threshold model accuracy based at least in part on the per-image quantitative image saliency metric. (Item 13) 13. The method of any one of items 1 or 2-12, further comprising maintaining a library of device-specific parameters and training data, and generating the updated device-specific ANN model is further based on the device-specific parameters and training data. (Item 14) Item 14. The method of any of items 1 or 2-13, further comprising distributing, by the processor, the updated device-specific ANN model to at least a subset of the user devices associated therewith. (Item 15) 1. A system for generating device-specific artificial neural network (ANN) models for distribution across user devices, the system comprising: one or more processors; a memory coupled to the one or more processors, the one or more processors executing a plurality of modules stored in the memory, the plurality of modules comprising: a user interface that receives instructions from a user to identify one or more sample data sets from the user device of a user environment, the sample data sets comprising media data and predictions from a device-specific ANN model executing on the user device; a data store containing the sample data set; a business logic module, the business logic module, when executed, (i) identifies use case datasets stored in the data store, the use case datasets comprising at least training data parameters, (ii) identifies training data from the sample datasets that satisfy the training data parameters provided in the use case datasets, and (iii) identifies a device-specific ANN model stored in the data store; an artificial intelligence machine learning module, which when executed generates an updated device-specific ANN model from each stored instance of the device ANN model based on the training data; and a memory; A system comprising: (Item 16) Item 16. The system of item 15, wherein the user device comprises a plurality of image capture devices. (Item 17) 17. The system of claim 16, wherein the media data comprises image data, and application of the ANN model to the image data facilitates identification of an object of interest within the image data. (Item 18) 18. The system of claim 15, 16, or 17, wherein the training data parameters include media data parameters and device parameters. (Item 19) Item 19. The system of item 18, wherein the media data parameters include one or more of a color index, a brightness index, a contrast index, an image temperature, a color tone, one or more hue values, and a gamma value. (Item 20) 20. The system of claim 18 or 19, wherein the device parameters include one or more of available memory, processing speed, image resolution, and capture frame rate. (Item 21) 21. The system of claim 15 or 16-20, wherein the use case dataset is specific to a particular use case. (Item 22) 22. The system of claim 21, wherein the use case comprises an environmental aspect and a functional aspect. (Item 23) 23. The system of claim 22, wherein the functional aspects of the use case include facial recognition. (Item 24) 24. The system of claim 22 or 23, wherein the environmental aspect of the use case comprises one of installing the device in an outdoor environment, installing the device in an indoor environment, installing the device in a well-lit environment, or installing the device in a poorly lit environment. (Item 25) 25. The system of claim 15, wherein the predictions generated by the device-specific ANN model comprise a quantitative image saliency metric for each image that indicates the likelihood that the media file contains an object of interest. (Item 26) 26. The system of claim 25, wherein the artificial intelligence machine learning module further determines the minimum number of images required to achieve a threshold model accuracy based, at least in part, on a quantitative image saliency metric for each image. (Item 27) The system of any of items 15 or 16-26, further comprising a library of device-specific parameters and training data, wherein the artificial intelligence machine learning module generates the updated device-specific ANN model based on the device-specific parameters and training data. (Item 28) The system of any of items 15 or 16-27, further comprising a deployment module for distributing the updated device-specific ANN model to at least a subset of the user devices associated therewith. (Item 29) 1. A method for optimizing the execution of a device-specific trained artificial neural network (ANN) model on an edge device, the method comprising: receiving, by a processor, a first trained ANN model and a second ANN model, wherein the first ANN model and the second ANN model each perform different estimations on input data, and an output of the first ANN model serves as an input to the second ANN model; merging the first ANN model, the second ANN model, and control flow execution instructions into a combined software package; deploying the combined software package to an edge device for execution thereon in accordance with the control flow instructions; and A method comprising: (Item 30) 30. The method of claim 29, wherein the first trained ANN model and the second trained ANN model each comprise respective analytical criteria and use case data, and the processor selects the first and second ANN models based, at least in part, on the analytical criteria therein. (Item 31) The method of claim 29 or 30, further comprising generating a parent ANN as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and delivering the meta-architecture to the edge device so that it executes as a single ANN model. (Item 32) 32. The method of claim 29, 30, or 31, wherein the edge device comprises a camera. (Item 33) Item 33. The method of item 32, wherein execution of the first ANN model and the second ANN model on the camera identifies an object of interest in an image file captured on the camera. (Item 34) 1. A system for optimizing the execution of a device-specific trained artificial neural network (ANN) model on an edge device, the system comprising: one or more processors; a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory, the computer-executable instructions, when executed, identifying, in a data store, a first trained ANN model and a second ANN model, wherein the first ANN model and the second ANN model each perform different estimations on input data, and an output of the first ANN model serves as an input to the second ANN model; merging the first ANN model, the second ANN model, and control flow execution instructions into a combined software package; deploying, by a distributed module, the combined software package to an edge device for execution thereon in accordance with the control flow instructions; memory and A system comprising: (Item 35) Item 35. The system of item 34, wherein the first trained ANN model and the second trained ANN model each comprise respective analytical criteria and use case data, and the processor selects the first and second ANN models based, at least in part, on the analytical criteria therein. (Item 36) The system described in item 34 or item 35, wherein execution of the instructions further generates a parent ANN as a meta-architecture based on the first ANN model architecture and the second ANN model architecture, and the meta-architecture is delivered to the edge device so that it executes as a single ANN model. (Item 37) 37. The system of claim 34, 35, or 36, wherein the edge device comprises a camera. (Item 38) Item 38. The system of item 37, wherein execution of the first and second ANN models on the camera identifies an object of interest in an image file captured on the camera. (Item 39) 1. A method for identifying an object of interest in an image file, the method comprising: receiving one or more image files, each image file potentially including an object of interest; applying a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values ​​that indicate the likelihood that a particular pixel is part of the object of interest; generating a three-dimensional saliency surface map having x-axis, y-axis, and z-axis, where the x-axis and y-axis values ​​define pixel locations within the image and the z-axis value is a saliency value specific to the pixel; selecting a curve shape from a library of curve shapes, applying the curve shape to the saliency surface map, and determining a fit between the curve shape and the three-dimensional surface; determining whether the image file contains the object of interest based on the match; A method comprising: (Item 40) Item 39. The method according to item 39, wherein the curve shape is selected based on the object of interest. (Item 41) 41. The method of claim 39 or 40, wherein the curve shape is selected from one of a Gaussian distribution, a Poisson distribution, and a hybrid distribution. (Item 42) 42. The method of claim 39, further comprising adding the image file to a library of image files for use in training an artificial neural network (ANN). (Item 43) 43. The method of claim 40, 41, or 42, wherein the ANN is trained to identify objects of interest within subsequent media files. (Item 44) Item 44. The method of any of items 40 or 41-43, wherein the ANN is trained to segment objects in subsequent media files. (Item 45) 1. A system for identifying an object of interest in an image file, the system comprising: one or more processors; a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory, the computer-executable instructions, when executed, receiving one or more image files, each image file potentially including an object of interest; applying a non-binary ground truth label to each image file, the non-binary ground truth label comprising a distribution of pixel-specific saliency values ​​that indicate the likelihood that a particular pixel is part of the object of interest; generating a three-dimensional saliency surface map having x-axis, y-axis, and z-axis, where the x-axis and y-axis values ​​define pixel locations within the image and the z-axis value is a saliency value specific to the pixel; selecting a curve shape from a library of curve shapes, applying the curve shape to the saliency surface map, and determining a fit between the curve shape and the three-dimensional surface; determining whether the image file contains the object of interest based on the match; memory and A system comprising: (Item 46) Item 46. The system of item 45, wherein the curve shape is selected based on the object of interest. (Item 47) 47. The system of claim 45 or 46, wherein the curve shape is selected from one of a Gaussian distribution, a Poisson distribution, and a hybrid distribution. (Item 48) Item 48. The system of item 45, item 46, or item 47, wherein execution of the instructions further adds the image file to a library of image files for use in training an artificial neural network (ANN). (Item 49) Item 49. The system of item 48, wherein the ANN is trained to identify objects of interest within subsequent media files. (Item 50) 50. The system of claim 48 or 49, wherein the ANN is trained to segment objects in subsequent media files. (Item 51) 1. A method for storing image data for transmission of video data, said method comprising: receiving video data at an edge device in one of a plurality of standard video data formats; extracting a plurality of image slices from the video data, the image slices comprising an image, a start index time and an end index time indicating a temporal location of the image slice within the video data, and region of interest parameters describing two-dimensional coordinates of a region of interest within the image; A method comprising: (Item 52) Item 52. The method of item 51, wherein the receiving of the video data and the extraction of the image slices are performed on an edge device. (Item 53) 53. The method of claim 52, further comprising using one or more artificial neural networks on the edge device to analyze the image slices and determine the region of interest and whether the region of interest includes an object of interest. (Item 54) 54. The method of claim 53, further comprising identifying each image slice as high resolution if the image slice includes an object of interest, and otherwise identifying the image slice as low resolution. (Item 55) 55. The method of claim 54, further comprising transmitting the high-resolution image slices to an artificial intelligence machine learning module for inclusion into an artificial neural network of a training data set specific to the edge device on which the image was captured. (Item 56) Item 51 or the method of any of items 52-55, wherein the standard video data format comprises an H.264 data format. (Item 57) 1. A system for storing image data for transmission of video data, the method comprising: one or more processors; a memory coupled to the one or more processors, the one or more processors executing computer-executable instructions stored in the memory, the computer-executable instructions, when executed, receiving video data at an edge device in one of a plurality of standard video data formats; extracting a plurality of image slices from the video data, the image slices comprising an image, a start index time and an end index time indicating a temporal location of the image slice within the video data, and region of interest parameters describing two-dimensional coordinates of a region of interest within the image; memory and A system comprising: (Item 58) Item 58. The system of item 57, wherein the receiving of the video data and the extraction of the image slices are performed on an edge device. (Item 59) 59. The system of claim 58, wherein execution of the computer-executable instructions further comprises using one or more artificial neural networks on the edge device to analyze the image slices and determine the region of interest and whether the region of interest includes an object of interest. (Item 60) 60. The system of claim 59, wherein execution of the computer-executable instructions further identifies each image slice as high resolution if the image slice includes an object of interest, and otherwise identifies the image slice as low resolution. (Item 61) 61. The system of claim 60, wherein execution of the computer-executable instructions further transmits the high-resolution image slices to an artificial intelligence machine learning module for inclusion into an artificial neural network of a training data set specific to the edge device on which the image was captured. [Brief explanation of the drawings]

[0030] In the drawings, like reference characters generally refer to the same parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the implementations. In the following description, various implementations are described with reference to the following drawings:

[0031] [Figure 1] FIG. 1 is an illustration of a general arrangement of components operating in an environment in which various embodiments of the present invention may be implemented.

[0032] [Figure 2] FIG. 2 illustrates an exemplary data architecture according to various embodiments of the present invention.

[0033] [Figure 3]FIG. 3 is a flowchart illustrating a model training process that may be implemented and performed according to various embodiments of the present invention.

[0034] [Figure 4] FIG. 4 is a flowchart illustrating an exemplary method for developing a training dataset according to various embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] Detailed Description Described herein, in one embodiment, is a method and support system for generating, deploying, and maintaining endpoint deployable artificial intelligent systems, machine learning mechanisms, and data models, implemented as a comprehensive platform. As shown in FIG. 1 , platform 100 implements a framework containing a front-end user interface (“user interface”) 105 for user interaction, a business logic module 110, a data store 115, an artificial intelligence / machine learning (“AI / ML”) training module 120, and deployment tools integrated into a user environment 125. The framework components may be communicatively coupled using one or more APIs (130a, 130b, 130c, and 130d) as provided by platform 100.

[0036] According to some embodiments, platform 100 provides users of the platform with one or more user interfaces 105 for accessing datasets 140, such as those provided by users and collected from endpoint devices, and otherwise performing data analysis thereon. These user interfaces 105 may be provided using, inter alia, distributed, localized applications (e.g., SDKs, APKs, IPAs, JVM files, other localized executables, and the like), APIs (e.g., JSON, REST, other data transfer protocols, and the like), websites, or collections of web application functionality, in combination or separately. User interfaces 105 facilitate the provision of analysis criteria to the criteria collection system. Analysis criteria may include configurations, parameters, and access to the user's datasets. According to some embodiments, configurations and parameters may be used as, or otherwise referred to as, use case data. Use cases may include functional processes such as facial recognition, license plate and other character recognition, image detection processes for identification authentication, object detection for autonomous driving applications, motion detection and intruder alerts, and others. Use cases may also include environmental aspects such as outdoor versus indoor installation, nighttime versus daytime, crowded spaces (e.g., airports, transit stations) versus sparsely populated spaces (bank security camera deployment, home camera deployment, etc.).

[0037] Importantly, the edge devices used within each use case may differ and often have device-specific characteristics and processing limitations that, in many embodiments of the present invention, are accounted for by and / or incorporated into the models used in those devices. Examples of device-specific characteristics include device-specific characteristics, such as available memory, processing speed, image resolution, capture frame rate, and others.

[0038] For example, the user interface 105 can provide analytical feedback to a user based on the uploaded dataset 140. The dataset provided by the user (referred to herein as a “media dataset”) may include, without limitation, a single image file, multiple image files, compound image files containing multiple images therein (e.g., GIF, APNG, WebP, among others), video files containing one or more frames, multiple video files, audio files, among others. The feedback may include data that characterizes the dataset before performing further operations on such dataset, such as image dataset classification and balancing. The feedback may be qualitative (e.g., high quality, low quality, etc.) or quantitative in nature, such as, for example, one or more quality metrics for the training dataset describing various photometric properties (lightness, luminance, color spectrum, etc.) and geometric properties (shape, edge definition, etc.) of the images and potential objects of interest within the images.

[0039] Datasets containing images may be further analyzed to extract or otherwise generate media properties for the images and other objects contained therein. Media properties may include, but are not limited to, color index, brightness index, contrast index, and other image properties (e.g., temperature, tone, hue, gamma), among others. Datasets containing more than one image, such as a composite image or video file, may be analyzed as a batch to identify, extract, or otherwise generate media properties for the multiple image or video files of the dataset.

[0040] The platform 100 also generates other media properties, such as a complexity index, from the media dataset provided by the user. The complexity index may be a set of diagnostic data representing the complexity of an image or one or more video frames of the media dataset. The platform may further compare media properties within the media dataset provided by the user, such as those associated with images, frames of video files, frames between video files, or between video files themselves. The user interface 105 of the platform 100 may also be used to identify or further generate comparisons of media properties or other characteristics of or between media datasets. For example, the platform may generate a comparison of the background and foreground of an image dataset, such as those found in an individual image or individual video frame. Similarly, the platform may also generate comparisons of objects of interest and other objects of no interest contained within the media dataset, such as distinguishing people in an image from background objects. Additionally, the platform may assign classes to the media datasets for further comparison therebetween. Examples of classes may include general categories such as a person, a human face, a car, an animal, or a defect in a manufactured product, or specific classes such as a person in a certain vicinity, a person within a certain distance, an adult German Shepherd, an adult Dalmatian, a baby Labrador, or a crack in a material, a contaminated material, or a chip on a material.

[0041] In other embodiments, the platform 100 can generate a quantitative image saliency metric for images in the image dataset, which may comprise a single number or a matrix of numbers or regions or other measurements assigned at the pixel level, which can be used to predict the difficulty with which a computational process can distinguish among objects of interest and / or between objects of interest and the background in an image. Based on the image saliency metric, the minimum number of images required to train a model to achieve a specified accuracy can be determined. The process can be extended using human-readable criteria such as brightness, contrast, distance from the camera, etc. to provide additional image collection recommendations to further improve and refine the training dataset. For example, the platform may identify that the training data contains a set of dark / distant images and dark / close images, but that adding brighter / distant images would result in a significantly improved training dataset. Similarly, if the training data contains high-quality images with significant contrast values, adding additional images to the training data may not be necessary or may only slightly increase the accuracy of the model.

[0042] According to some embodiments, the user interface 105 can provide recommendations to the user based on the feedback associated with that media dataset 140. In some examples, the recommendations may be provided along with or otherwise included in the feedback. Recommendations as provided by the platform may include, without limitation, suggestions for additional data for the user to collect and include in the media dataset, as well as suggested extensions to one or more media datasets for applying improvements thereto.

[0043] According to some embodiments, the analysis performed by platform 100 may be implemented by machine learning mechanisms or artificial neural networks (“ANNs”). To implement such analysis, the platform may further include a criteria collection system and provide users with access to artificial intelligence tools and the ability to use a front-end user interface. For example, one or more user interfaces may be provided to collect key analysis criteria from the user regarding requirements or other preferences, whereby the platform may use to analyze the user's media dataset. For example, the user may identify analysis criteria including, but not limited to, speed and latency requirements as required by the user's implementation, hardware and network requirements as required by the user's implementation, size of objects to be identified within the media dataset, reaction time tolerance as required by the user's implementation, tolerance for false positives as identified by the platform, tolerance for non-detections as identified by the platform, and accuracy requirements for predictions made by the platform, among others. In some instances, the criteria collection system may also facilitate filtering of large datasets down to datasets that meet certain image criteria or size limits.

[0044] According to some embodiments, the platform uses its intelligent systems (e.g., machine learning mechanisms, artificial neural networks, and the like) to identify key analytical criteria that are best suited for the user's implementation. Some embodiments of the platform's criteria collection system use dual (or multiple) ANNs to provide the user with access to the best artificial intelligence tools and capabilities for their associated use cases. In other words, a first neural network may receive a media dataset as provided by the user and determine the best analytical criteria to be used by a second neural network to perform a particular analysis on the same or other media dataset as provided by the user. For example, a user may upload a video clip of a sample use case to the first neural network. The user may identify common use cases or objects, whether selected from a list or identified in a custom manner by the user, to be used for video clip analysis. Based on the user's selection, the first neural network analyzes the video clip provided by the user and determines the analytical criteria necessary for the second neural network to further appropriately analyze the uploaded video clip. For example, a first ANN may be used to identify regions of interest in an image that are likely to contain a person within an image with multiple other objects, while a second ANN may be used to process the regions of interest and perform facial recognition on the image of the person. In some cases, the analysis criteria may be automatically extracted from a video clip and may include a reaction time, a specific definition of an accuracy metric, and a quantitative value of the metric. The first neural network may determine the required "reaction time" of the second neural network, the size of the object to be identified by the second neural network, or even the ideal number of frames of video that the second neural network can use at runtime to correctly determine a "response."

[0045] According to some embodiments, the platform may further include intelligent operation tools to facilitate the implementation and maintenance of a user's intelligent systems (e.g., machine learning mechanisms, artificial neural networks, and the like). For example, the platform may provide the user with an integrated compilation of software applications or software development kits (SDKs) for the user's specific target hardware. The SDK compilation may contain a unique license (e.g., token) embedded therein or associated therewith, however, other licensing models may also be used. The software (e.g., SDKs, other software applications, etc.) facilitates monitoring of performance and / or statistical information about the hardware on which the software runs, as well as the software and communications between various platform components. The software may further provide a comparison of statistical information about the horizon to statistical information about the training data.

[0046] In some embodiments, the platform uses data obtained by software distributed across the user's hardware to provide training data and recommendations for configuring the intelligent system to the user. The platform may also provide a media dataset as collected by the user's hardware at runtime with predictions superimposed thereon. In doing so, the platform may further provide a user interface for the user to mark predictions provided in the runtime data as correct, incorrect, or, in some cases, grade them along a gradation of accuracy (e.g., a numerical value, probability, qualitative tag, etc., representing the likelihood that the prediction is correct) to facilitate a semi-supervised learning environment. In response to receiving an indication that the prediction is correct, the platform may add the associated runtime data to an auxiliary training dataset. Adding the runtime data with the correct or corrected prediction to the auxiliary training dataset facilitates the continuous training of a semi-supervised machine learning procedure that updates an ANN model (or other artificial intelligent model) for use by the user's intelligent system. Once updated, the ANN model is deployed to the user's hardware, distributing improvements to the user's intelligent system.

[0047] The AI / ML training system accesses the user's dataset as provided by the user according to the configuration and parameters of the analysis criteria and generates a subsample of training data. For example, the configuration and parameters as provided by the analysis criteria may include a request to limit the training data to datasets with faces close to the camera and exclude faces that are further away. According to some embodiments, generating the training data can be expanded or further defined based on the type of device collecting the media dataset. Device type data can be implemented using adaptive diffusion, as described below.

[0048] Once appropriate training data is collected, the AI / ML training system generates a new ANN model and trains it according to the analysis criteria. The platform's AI / ML training system may store the trained ANN model and other models in a data store for retrieval when required. Storing the trained model may further include storing associated training metadata and associated analysis criteria (e.g., configurations and parameters), both of which may be included as use case data. According to some embodiments, the use case data may indicate how a particular model may be used and / or what such a model may be intended for. For example, a model may be used to implement selective attention on a media dataset or even to extract areas therein.

[0049] According to some embodiments, the AI / ML training system searches the data store for trained models with a meta-architecture that can best implement or otherwise handle the data indicated by the use case data. Thus, the data store may be searched or otherwise filtered based on the use case data (e.g., analytical criteria, training metadata) of the stored models. According to some embodiments, similar use case data across multiple models may indicate the meta-architecture of the models stored therein.

[0050] For example, an AI / ML training system may search its associated data store for ANN models trained to detect objects of interest of a particular size. A model identified by this search may then be defined as a particular meta-architecture, representing an architecture capable of identifying objects of interest at a particular size. Similarly, an AI / ML training system may search its associated data store for ANN models trained to analyze the relative complexity of the foreground and background of a media dataset received as input. A model identified by this search may then be defined as a particular meta-architecture, representing an architecture capable of analyzing the relative complexity of the foreground and background of a media dataset.

[0051] According to some embodiments, the meta-architecture may be further identified or otherwise organized as a custom meta-architecture within the data store. A custom meta-architecture may be identified by a use case for the underlying model, such as a model used for selective attention or a model used for object detection. According to some embodiments, the ANN itself, as well as other trained search models, may be used to perform or otherwise extract results from a data store associated with an AI / ML training system. Thus, one or more searching ANNs may be used to identify candidate meta-architectures that contain models (or otherwise models themselves) similar to those of a use case identified by a user. For example, a user may provide the search ANN with analysis criteria or other data indicative of a model for determining the complexity of a media dataset, and the search ANN may then return a meta-architecture (or otherwise models therein) indicative of such a use case.

[0052] According to some embodiments, the meta-architecture search used for the search ANN may be similarly trained according to other ANNs provided by the platform. The search ANN may also be trained according to a unique loss function. For example, the search ANN may be trained using a selective attention metric, among other techniques. Furthermore, the search ANN may be optimized according to various characteristics required by a particular search, such as a particular search order, priority, density, and depth of the search space, among others. Similarly, the search ANN may be optimized according to a Bayesian optimization strategy, Gaussian process, or otherwise using statistical weighting to determine correlations between analytical criteria (e.g., training cycle parameters) and data associated with the training data and / or use case data.

[0053] An AI / ML training system may use an ANN to find an optimal error threshold for a particular model and use case according to analytical criteria, such as those provided by a user, among other data. For example, a model with a use case for finding a cluster of pixels (or region of interest or “ROI”) representing an object of interest based on a three-dimensional map of inputs (e.g., x location, y location, and probability that the object of interest is present at that location) may be given a particular error threshold. Thus, the ANN may determine a higher error threshold for a model with a similar use case with an additional level of complexity, such as another input dimension (e.g., x location, y location, probability that the object of interest is present at that location, and a time index of a particular frame). One approach for identifying regions of interest and objects within those regions is described in U.S. Patent Application Serial No. 16 / 953,585, the entire disclosure of which is incorporated herein by reference.

[0054] In one embodiment of the present invention, a ground truth polygon mask (or "ground truth labels") may be used to define a ROI within an image. In conventional techniques, a binary decision is made based on pixel values, such that pixels inside the polygon are considered part of the object, while pixels outside the polygon are considered "not the object." In one embodiment of the present invention, a "pixel saliency value" can be assigned as a z-value for each xy pixel location within the ground truth polygon that represents the likelihood that the pixel is part of the object, and a saliency surface map can be generated from the ROI. In some cases, pixels or groups of pixels that meet a certain likelihood threshold can be inferred to be part of the object.

[0055] In some cases, instead of (or in addition to) independently calculating or assigning a saliency value to each pixel, a curve shape can be applied to the saliency surface map based on the expected object within the ROI, such as a head shape if a human face is expected. A curve shape associated with "head" (e.g., a hat) can be used to make an inference as to whether the object is a head. In some instances, each pixel may be assigned an initial value based on a predetermined distribution for that object, and a difference value may be calculated. For example, face recognition may be best predicted using a "hybrid Gaussian" curve, where an initial gradual increase in saliency occurs at the edge of the ROI, and values ​​across the ROI follow a Gaussian gradient shape, such that pixels closer to the center of the ROI have higher saliency values ​​than those along the edge. In some cases, different curve shapes may be used to infer the presence of different objects of interest within the ROI. For example, for smaller, persistent objects such as road signs, a Poisson distribution may be used to assign saliency values ​​to pixels, while a different distribution may be used for larger objects where edge boundaries are important, such as cars or other vehicles. The "fit" between a particular shape (or set of shapes) and an object of interest can then be used to further train an object ANN model for subsequent object detection.

[0056] These gradient values ​​can be applied to various images and, based on fitness and accuracy, used as input into a training step to further refine each model for a specific object use case, device, or combination thereof.

[0057] According to some embodiments, once a model has been identified by an AI / ML training system, it may be further trained or otherwise optimized, for example, using transfer learning (e.g., adaptive diffusion), as described below.

[0058] According to some embodiments, the AI / ML training system may further use and merge two or more models, in some cases, across remote hardware in the customer environment to determine model collaboration structures and facilitate intelligent distribution of data. For example, a model used for selective attention may be structured or otherwise organized so that it does not operate on the image data it initially receives, but instead forwards feature maps to a second model, such as an object detection model. This deployment option may be useful, for example, in instances where an initial selective attention model and a second object detection model are combined into a single wrapper function and deployed to an endpoint device provided via control flow software. In such examples, a “switched pipeline” implementation may be used, where two (or in some cases, more than two) models can be run on the same input data, either in parallel or sequentially, as directed by a user or a pre-configured switch. Thus, a device that may only support the execution of a single model due to processor and / or power constraints can use two distinct but “merged” ANN models to perform two different inferences (e.g., find an area of ​​interest in an image, then an object within the area of ​​interest).

[0059] The model collaboration structure as determined by the AI / ML training system may be associated with the business logic component of the platform, stored by a data store, or otherwise represented.

[0060] According to some embodiments, the platform may build SDKs and other software from a single ANN model, such as provided by an AI / ML training system. Thus, a compiler implemented by the platform can compile models for multiple different hardware architecture targets for execution on specific hardware located in the customer environment. In traditional implementations, ANN models are trained using hardware-independent parameters. This approach simplifies training and deployment, but suffers from accuracy and performance issues. By doing so, slight changes to the processing may be introduced once the model is compiled, which may result in suboptimal processing on a given hardware device. To address this issue, in some embodiments of the present invention, the training process is performed on a processor (or emulator) specific to the edge device hardware (e.g., a particular model of camera). Analyzing the results of the training step on specific hardware allows the model to be trained for that specific device, resulting in a hardware-specific variant of the model optimized for the processor used within that device. In some cases, a “library” of emulators is provided for processing training data and models for each specific device.

[0061] According to some embodiments, the SDK and other software are distributed using a unique DRM system. In one example, the unique DRM system provides a heartbeat-like system in which data can be transmitted to a central licensing and authorization server based on a predetermined time period. The data transmitted during each beat of the heartbeat-like system may include data and metadata for the media dataset and images themselves, such as, but not limited to, an indication of the scene, location (e.g., GPS coordinates), detected event data such as width, height, and time index, estimated speed, frequency of SDK usage, and other data associated therewith. The predetermined time period may be unique to each user environment or each device. In some examples, a risk engine may be incorporated into the unique DRM system to determine whether a license associated with a customer environment utilizes a longer or shorter predetermined time period. The risk engine may also determine whether a license should be denied based on one or more license restrictions, usage limits, time periods, etc. Additionally, data received by the DRM system from devices or other endpoints that are deemed suspicious or otherwise suspect may be brought to the user's attention.

[0062] According to some embodiments, the unique DRM system may also track how software and other data is used by hardware in the customer environment. For example, the DRM system may track, among other data, the use of specific models, the detections associated with each model, and the use case detections. Thus, the DRM system can determine prices for users based on the software and data usage tracked on each device or hardware in the user environment.

[0063] According to some embodiments, the data provided within the heart rate-like system may be used to identify faulty devices or devices that have been tampered with or jammed through a heuristic or risk engine type system.

[0064] The platform further provides the Visual Intelligence SDK with various artificial intelligence (AI) detection mechanisms. According to some embodiments, the Visual Intelligence SDK may include features such as inference, image processing, unique DRM, quality control sampling, and over-the-air updates, among others. The Visual Intelligence SDK may implement a dynamic post-processing analysis engine to detect shared terms across one or more sequential images from one or more devices across a user environment.

[0065] Similarly, the Visual Intelligence SDK's shared term detection may be similarly implemented across one or more user environments to detect shared terms across one or more user environments. In such cases, the process crosses multiple frames of an image to increase the number of frames to be classified and the voting strategy, prior to edge unfolding, so that inconsistencies can be mitigated, thus improving model accuracy. More specifically, a user can define a temporal window of frames (e.g., 10 frames) from a video file to perform multiple estimations on some or all frames within that window. A comparison is performed to measure the difference between the image in each frame and an image known to contain the object of interest (e.g., a particular person's face). If a certain percentage of frame embedding (e.g., 50%, which can be user-defined) is below a specified distance threshold (again, which can also be user-defined), the object is considered to be identical to the object captured from multiple frames.

[0066] The visual intelligence SDK may also implement privacy features such as those provided by the platform's privacy ANN or other artificial intelligence model. For example, a model trained for selective attention may be implemented to filter out sensitive information (e.g., faces, PPI, nudity, or other sensitive data) for use in quality assurance (QA) sampling by the platform. Similarly, a model trained for selective attention may be distributed in edge devices or other hardware within the user environment to filter out sensitive information before further analysis or transmission to other devices. The model may accomplish such privacy filtering by encrypting, redacting, obfuscating, or compressing the field of view of a media dataset, or otherwise using cropping features to remove sensitive data. The model may also extract specific areas (e.g., objects of interest) from a media dataset (removing the rest of the dataset) for privacy purposes. For example, after identifying an object of interest in a media dataset, one or more models may extract only the object of interest and remove the rest of the environment in the field of view, maintaining the privacy of others in the rest of the environment. Therefore, the smaller images extracted by the model also contain data annotations therein, identifying the location of the extracted image within the original field of view and facilitating the construction of an original media dataset from which non-essential data has been removed.

[0067] In some embodiments, privacy filtering may encrypt a field of view or a portion of a field of view using a technique that requires multiple factors for decryption. In some embodiments, one such factor is a unique token that changes over time. In such embodiments, a media image or video may be decrypted by using only user authorization factors in conjunction with a specific token that corresponds to a certain period of time on a specific device or group of devices. In some embodiments, the decryption token is stored digitally so that readings of the token are recorded, for example, for audit or control purposes.

[0068] Unique data structure (for transmission of video data with variable resolution) Similar to the privacy features implemented by the visual intelligence SDK detailed above, the visual intelligence SDK (or other software provided by the platform) may also extract specific areas (e.g., objects of interest) from a media dataset (removing the rest of the dataset) to generate a media dataset of smaller images and reduce file size for transmission. Thus, the smaller images extracted by the model may further include data annotations therein identifying the location of the extracted image within the original field of view and facilitating the construction of an original media dataset from which uninteresting data has been removed. By removing uninteresting data from the media dataset before transmission, the smaller media dataset can be transmitted over a network at a significantly reduced file size. Referring to FIG. 2 , the data file architecture (205), which represents video data using the H.264 protocol, includes a significant amount of data not required for image detection and extraction according to the techniques described herein. Instead, a slice (or slices) (210) of data is selected to include content encapsulating one or more regions of interest from a video segment. The contents of a "slice" 210 may include various parameters 215 related to a region of interest within an image. The parameters may include, for example, time index start and end, top, bottom, right, and left coordinates, and the extracted image itself, or in some cases, a downsampled version of the image.

[0069] The reduced file size may further be accomplished at an edge device in the user environment by transmitting areas of the media dataset containing the object of interest at high resolution while transmitting the remainder of the field of view at low resolution. The high-resolution and low-resolution areas of the digital media dataset may be transmitted as a reduced media dataset with a file size significantly smaller than the original media dataset. Similar to the reconstruction described above, the reduced media dataset may further include data annotations therein identifying the locations of the high-resolution and low-resolution images within the original field of view and facilitating the construction of an original media dataset with areas containing only the object of interest in high resolution. The video file may repeat the file reduction and construction process for each frame of the video file.

[0070] The AI / ML training system of a platform as described herein may further provide the ability for pre-trained models to be further trained using data associated with real-time data collected by edge devices and other hardware in the user's environment. As described above, once a model is identified by the AI / ML training system, it may be further trained or otherwise optimized using a training method called adaptive diffusion (e.g., continuous transfer learning methods). Alternatively, adaptive diffusion may be performed on an ANN model that has already been distributed to hardware in the user's environment, for example, by federated learning techniques.

[0071] According to some embodiments, and referring to FIG. 3 , the user's environment may contain an ANN model based on or otherwise sourced from the data store of the AI / ML training system 120 of the platform 100. Quality assurance (QA) sampling data (“training data”) may be collected from the model and transmitted to a review module (step 305) and used as an initial training set for the initial ANN / AI model (step 310). The model may then be deployed, and results may be collected from its use in the field (step 315). The QA sampling data may be reviewed (step 320) using an automated review process and / or human review (e.g., reviewed by a user). The automated review process may be performed by the ANN model to identify relevant (e.g., correct) data from the QA sampling data. Alternatively, a human (e.g., a user) may review the QA sampling data for accuracy and manually mark results as either correct or incorrect. Similarly, a human may review the QA data in conjunction with the ANN model to prompt human review of the QA sampling data. The reviewed QA sampling data (e.g., data marked as correct) may be stored as training data or otherwise incorporated into training data stored in a data store to apply updates to the ANN model during retraining. Once the training data is updated, an updated ANN model may be generated or otherwise trained therefrom (step 325). The platform may distribute the updated ANN model to hardware in the user's environment to provide a better trained model for the user's particular use case.

[0072] According to some embodiments, and with reference to FIG. 4 , an adaptive diffusion procedure as described above may be provided or otherwise implemented differently for different hardware devices in a user environment. For example, QA sample data may be reviewed by an ANN model and, after review, provided to a data store where it is associated with the particular device's individual training data. Similar to the training described above, the updated training data may be reviewed by an ANN model, a human, or a combination thereof. The reviewed QA sample data stored in the data store may be correlated with feedback data or other media information associated with the media dataset, such as brightness, background complexity, size or geometry of objects of interest, and data associated with the device (e.g., statistical device information), among other things. The ANN model may select the reviewed QA sample data and retrain or further update the ANN model for one or more target devices. Thus, each particular device may receive updated or further updated model training using only the reviewed QA sample data received from that particular device. Alternatively, the QA sample data may be provided in a data store associated with training data for a particular group of devices, allowing the group of devices to receive training and further updated models using only the scrutinized QA sample data received from that particular group of devices.

[0073] More specifically, QA sample data is collected and / or received from various sources representing a specific use case (step 405), and an initial AI model is trained using the collected data (step 410). As with the previous use case, the AI ​​model is deployed into a field corresponding to the use case, and results are collected (step 415). In this embodiment, the results can be collected and stored separately or otherwise identified as originating from different devices or deployments (step 420), while in other instances, the data can be combined into a single dataset. The results are then reviewed for accuracy using either an automated process, human review, or a combination of both (step 425). Once deemed accurate and sufficient, the images and associated results are corrected and used to create an updated training dataset (step 430); in some instances, the dataset can be segmented so that specific datasets are assigned to specific devices or groups of devices (step 440). The grouping can be based on some commonality, such as the camera manufacturer, make, and / or model, the environment in which the devices are used (e.g., training sets for images installed outdoors versus indoors, nighttime images versus daytime images, etc.), and / or functional use case commonality, such as facial recognition versus character recognition. Once updated, training datasets are created for the specific devices, which are then used to train updated AI models for the specific devices (step 450). The process then iterates over time as the new AI models are deployed and used in the field, new data is collected, and the process is repeated, resulting in continually improved AI models specific to the particular device and / or use case.

[0074] In addition to the training described above, the platform may further provide a retraining maintenance module that monitors models for all devices in the user's environment and automatically retrains each one. The retraining maintenance module may incorporate an ANN or other model and determine monitoring criteria associated therewith. For example, the retraining maintenance module may use Bayesian optimization to determine how often the retraining maintenance module should retrieve QA sample data for quality assurance and quality control purposes. Furthermore, the retraining maintenance module may further use the ANN or other model to determine how devices in the user's environment should be grouped or otherwise organized for training. For example, the retraining maintenance module may optimize the size of each device group based on its use case or other data associated with the devices. Devices may share curated QA sample data (e.g., auxiliary training data) for each device group to most effectively train each ANN model therein.

[0075] Implementations of the subject matter and operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and structural equivalents thereof, or in a combination of one or more of them. Implementations of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by or to control the operation of a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated, propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a receiver apparatus suitable for execution by the data processing apparatus.

[0076] A computer storage medium can be or be contained within a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. Further, a computer storage medium is not a propagated signal, although a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be or be contained within one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0077] The operations described herein may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0078] The term "data processing apparatus" encompasses all types of apparatus, devices, and machines for data processing, including, by way of example, a programmable processor, a computer, a system on a chip, or a plurality or combination of the foregoing. An apparatus may include special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, an apparatus may also include code that creates an execution environment for the computer program, such as code comprising processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of these. The apparatus and execution environment may implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0079] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored within a portion of a file that holds other programs or data (e.g., one or more scripts stored in markup language resources), within a single file dedicated to the program, or within multiple cooperating files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer or on multiple computers, located at one site or distributed across several sites and interconnected by a communications network.

[0080] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows may also be performed by, and an apparatus may also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0081] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer include a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or be operatively coupled to receive data from, transfer data to, or both. However, a computer need not have such devices. Furthermore, a computer can be embedded within another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name a few. Suitable devices for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0082] The phrases and terminology used herein are for the purpose of description and should not be regarded as limiting. The indefinite articles "a" and "an," as used in the specification and claims, should be understood to mean "at least one," unless clearly indicated otherwise. The word "and / or," as used in the specification and claims, should be understood to mean "one or both" of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with "and / or" should be construed in the same manner, i.e., as "one or more" of the elements so conjoined. Other elements than those specifically identified by the "and / or" clause may optionally be present, whether related to those elements specifically identified or not. Thus, as a non-limiting example, a reference to "A and / or B," when used in conjunction with open-ended language such as "comprising," may, in one embodiment, refer to only A (optionally including elements other than B), in another embodiment, refer to only B (optionally including elements other than A), in yet another embodiment, refer to both A and B (optionally including other elements), etc.

[0083] As used in this specification and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be construed as inclusive, i.e., the inclusion of at least one (but including more than one) of the elements or list, and optionally additional unlisted items. Only terms clearly indicated otherwise, such as "only one of" or "exactly one of," or, when used in the claims, "consisting of," will refer to the inclusion of exactly one element of the elements or list. Generally, the term "or," as used, shall only be construed to indicate exclusive alternatives (i.e., "one or the other, but not both") when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of." When used in the claims, "consisting essentially of" shall have its ordinary meaning as used in the field of patent law.

[0084] As used in this specification and claims, the phrase "at least one," in reference to a list of one or more elements, should be understood to mean at least one element selected from one or more of the elements in the list of elements, but does not necessarily include at least one of every element specifically listed in the list of elements, and does not exclude any combination of elements in the list of elements. This definition also allows for the optional presence of elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, whether related to those elements specifically identified. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B" or, equivalently, "at least one of A and / or B") can refer in one embodiment to at least one A, optionally including more than one A (and optionally including elements other than B), in the absence of any B; in another embodiment to at least one B, optionally including more than one B (and optionally including elements other than A); in yet another embodiment to at least one A, optionally including more than one A, and at least one B, optionally including more than one B (and optionally including other elements); etc.

[0085] The use of "including," "comprising," "having," "containing," "involving," and variations thereof, is meant to encompass the items listed thereafter and additional items.

[0086] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify claim elements does not, by itself, imply any priority, precedence, or ordering of one claim element relative to another, or the chronological order in which acts of a method are performed. Ordinal terms merely distinguish one claim element having a certain name from another element having the same name (but for the use of the ordinal term), and are used as labels to distinguish between claim elements.

[0087] Features that are described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, various features that are, for brevity, described in the context of a single embodiment may also be provided separately or in any suitable subcombination. Applicant hereby notifies that new claims may be formulated to such features and / or combinations of such features during patent prosecution of this application or any further application derived therefrom. The described device and system features may be incorporated into / used in corresponding methods, and vice versa.

Claims

1. 1. A method for optimizing the execution of a device-specific trained artificial neural network (ANN) model on a device, the method comprising: a processor receiving a first trained ANN model and a second trained ANN model, each of the first trained ANN model and the second trained ANN model performing different inferences on input data, the first trained ANN model and the second trained ANN model being trained using a same training data set; merging the first trained ANN model and the second trained ANN model and control flow execution instructions into a combined software package; deploying the combined software package to the edge device to execute the combined software package on the edge device in accordance with the control flow execution instructions; A method comprising:

2. The method of claim 1 , wherein the first trained ANN model and the second trained ANN model each comprise distinct analysis criteria and use case data.

3. The method of claim 1 , wherein the processor selects the first trained ANN model and the second trained ANN model based at least in part on analytical criteria.

4. 2. The method of claim 1, further comprising generating a parent ANN as a meta-architecture based on an architecture of the first trained ANN model and an architecture of the second trained ANN model, wherein the meta-architecture is delivered to the edge device such that the edge device executes it as a single ANN model.

5. The method of claim 4 , wherein the edge device comprises a camera.

6. The method of claim 5 , wherein execution of the first trained ANN model and the second trained ANN model on the camera identifies an object of interest in an image file captured on the camera.

7. 1. A system for optimizing the execution of a device-specific trained artificial neural network (ANN) model on an edge device, the system comprising: one or more processors; a memory coupled to the one or more processors; Equipped with The one or more processors execute computer-executable instructions stored in the memory, which, when executed, perform the following steps: identifying a first trained ANN model and a second trained ANN model, each of the first trained ANN model and the second trained ANN model performing different inferences on input data, and the first trained ANN model and the second trained ANN model being trained using an identical training dataset; merging the first trained ANN model and the second trained ANN model and control flow execution instructions into a combined software package; a distribution module deploying the combined software package to the edge device for executing the combined software package on the edge device in accordance with the control flow execution instructions; The system.

8. The system of claim 7 , wherein the first trained ANN model and the second trained ANN model each comprise distinct analysis criteria and use case data.

9. The system of claim 7 , wherein the processor selects the first trained ANN model and the second trained ANN model based at least in part on analytical criteria.

10. 8. The system of claim 7, wherein execution of the control flow execution instructions further generates a parent ANN as a meta-architecture based on an architecture of the first trained ANN model and an architecture of the second trained ANN model, and the meta-architecture is delivered to the edge device so that the edge device executes it as a single ANN model.

11. The system of claim 7 , wherein the edge device comprises a camera.

12. 12. The system of claim 11, wherein execution of the first trained ANN model and the second trained ANN model on the camera identifies an object of interest in an image file captured on the camera.

13. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method; The method comprises: a processor receiving a first trained ANN model and a second trained ANN model, each of the first trained ANN model and the second trained ANN model performing different inferences on input data, the first trained ANN model and the second trained ANN model being trained using a same training data set; merging the first trained ANN model and the second trained ANN model and control flow execution instructions into a combined software package; deploying the combined software package to the edge device to execute the combined software package on the edge device in accordance with the control flow execution instructions; 1. A non-transitory computer-readable storage medium comprising:

Citation Information

Patent Citations

  • Object detection device and object detection method and program

    JP2019008460A

  • Unsupervised data segmentation

    US20050147297A1

  • Visual object recognition

    US20170286809A1

  • Face recognition using stage-wise mini batching to improve cache utilization

    US20180060240A1

  • Server device, trained model providing program, trained model providing method, and trained model providing system

    WO2018173121A1