Object analysis

By combining hierarchical signal structure and object analyzer, the problem of low object analysis efficiency in images and videos with different resolutions and compression levels is solved, achieving more efficient object detection and recognition, and improving the computer vision function of autonomous vehicles.

CN114026608BActive Publication Date: 2025-10-28V NOVA INT LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080028395.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-13
Filing Date
2020-02-11
Publication Date
2025-10-28
Estimated Expiration
2040-02-11

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle images and videos with varying resolutions and compression levels in object analysis, especially in autonomous vehicles, leading to resource waste and low analysis efficiency.

Method used

By employing a hierarchical signal structure method, object detection and identification are achieved by generating and decoding signals at different quality levels, combined with an object analyzer that partially decodes the signal within the region of interest.

Benefits of technology

It improves the efficiency and accuracy of object analysis, reduces resource consumption, and enhances the performance of computer vision functions, especially in autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114026608B_ABST
    Figure CN114026608B_ABST
Patent Text Reader

Abstract

A method for performing object detection within a set of representations of a hierarchical structured signal, the set of representations including at least a first representation of the signal at a first quality level and a second representation of the signal at a higher second quality level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure involves object analysis. Background Technology

[0002] Object analysis can be performed on images and / or videos. Examples of object analysis types include, but are not limited to, object detection and object recognition. Object analysis can be performed on images and videos at different resolutions and compression levels, such as uncompressed or compressed image file formats. Examples of uncompressed file formats include BMP and TGA file formats. Examples of compressed image file formats include JPEG. Examples of compressed video file formats include H.264 / MPEG-4. Summary of the Invention

[0003] According to a first embodiment, a method is provided that includes performing object detection within a set of representations of a hierarchical structure signal, the set of representations including at least a first representation of the signal at a first quality level and a second representation of the signal at a higher second quality level.

[0004] According to a second embodiment, a method is provided that includes performing object analysis at least partially using a representation of a signal at a first quality level, wherein the representation of the signal at the first quality level is generated using a representation of the signal at a higher second quality level, wherein performing object analysis includes performing object detection and / or object identification.

[0005] According to a third embodiment, a method is provided that includes performing object analysis within an image in a multi-resolution image format, wherein multiple versions of the image at different corresponding image resolutions are available.

[0006] According to a fourth embodiment, a method is provided that includes partially decoding the representation of a signal in response to an object analysis performed within the representation of the signal detecting an object in a region of interest within the representation of the signal, wherein the partial decoding is performed relative to the region of interest.

[0007] According to a fifth embodiment, an apparatus is provided configured to perform the method described according to any one of the first to fourth embodiments.

[0008] According to a sixth embodiment, a computer program is provided that, when executed, performs the method described in any one of the first to fourth embodiments.

[0009] Other features and advantages will become clear from the following description, which is given by way of example only and with reference to the accompanying drawings. Attached Figure Description

[0010] Figure 1 A block diagram illustrating an example of a hierarchical system according to an embodiment;

[0011] Figure 2 A block diagram illustrating another example of a hierarchical system according to an embodiment;

[0012] Figure 3 A block diagram illustrating another example of a hierarchical system according to an embodiment;

[0013] Figure 4 A block diagram illustrating a portion of another instance of a hierarchical system according to an embodiment; and

[0014] Figure 5 A block diagram illustrating an example of a device according to an embodiment. Detailed Implementation

[0015] refer to Figure 1 An example of system 100 is shown. System 100 may include a distributed system. System 100 may be in an autonomous vehicle (also referred to as an "autonomous vehicle"). System 100 may be used to provide computer vision capabilities regarding autonomous vehicles.

[0016] In this example, system 100 includes a first device 110. In this example, the first device 110 includes an encoder 110. In this example, the encoder 110 generates encoded data. In this example, the encoder 110 receives data and encodes the received data to generate encoded data based on the received data.

[0017] In this example, system 100 includes a second device 120. In this example, the second device 120 includes a decoder 120. In this example, the decoder 120 generates decoded data. In this example, the decoder 120 receives encoded data from encoder 110 and decodes the encoded data to generate decoded data.

[0018] In this example, the second device 120 obtains data by receiving data from the first device 110. In some instances, the second device 120 obtains data in another manner. For example, the second device 120 may retrieve data from a memory, as will be referred to below. Figure 2 To describe in more detail.

[0019] The first device 110 and the second device 120 can be embodied in hardware and / or software. The first device 110 and the second device 120 can have a client-server relationship. For example, the first device 110 can act as a server, and the second device 120 can act as a client.

[0020] In this example, the first device 110 is communicatively coupled to the second device 120. Alternatively, the first device 110 may be directly communicatively coupled to the second device 120. A communication protocol can be defined for communication between the first device 110 and the second device 120.

[0021] In some instances, the second device 120 transmits data to the first device 110. For example, the second device 120 may transmit control data (also referred to as "feedback data") to the first device 110. Such control data may control how the first device 110 processes (e.g., encodes) data and / or control one or more other operations of the first device 110.

[0022] In this example, system 100 includes an object analyzer 130. The object analyzer 130 can be embodied in hardware and / or software. In this example, the object analyzer 130 is communicatively coupled to a second device 120. In this example, the object analyzer 130 is configured to perform object analysis within data being processed by the second device 120, as described in more detail below. The object analyzer 130 can control how the second device 120 processes (e.g., decodes) the data. Such control may include whether the second device 120 performs full decoding or partial decoding, and may include positioning (also referred to as a “region of interest (RoI)” or “bounded box”) to which other object analysis and / or decoding should focus, as described in more detail below. Such control may be based on the object analysis performed by the object analyzer 130, or on others.

[0023] Although in this example, the second device 120 and the object analyzer 130 are... Figure 1 The second device 120 and the object analyzer 130 are depicted as separate elements of system 100, but this separation can be logical rather than physical. Thus, although in some instances the second device 120 and the object analyzer 130 are provided as physically independent components of system 100, in other instances the second device 120 and the object analyzer 130 may be provided as a physical component of system 100.

[0024] System 100 may include Figure 1 One or more additional elements not shown in the diagram. Furthermore, although... Figure 1 The system 100 depicted includes a single first device 110 (e.g., including a single encoder), a single second device 120 (e.g., including a single decoder), and a single object analyzer 130, but in other instances, the system 100 includes more than one first device 110, more than one second device 120, and / or more than one object analyzer 130.

[0025] As noted above, in some instances, the first device 110 encodes the data and provides the encoded data to the second device 120, which then decodes the encoded data. In other instances, the first device 110 does not encode the data provided to the second device 120. In even more such instances, the data provided from the first device 110 to the second device 120 is not in an encoded form.

[0026] refer to Figure 2 An example of system 200 is shown. Figure 2 The instance system 200 described in the above reference includes Figure 1 The example system 100 described contains several elements that are the same as or similar to their counterparts. These elements are indicated using the same reference numerals, but with the addition of 100.

[0027] In this example, system 200 includes shared memory 240. Shared memory 240 may include double data rate (DDR) memory. In this example, shared memory 240 is shared between first device 210 and second device 220. In this example, first device 210 is communicatively coupled to shared memory 240. In this example, second device 220 is communicatively coupled to shared memory 240. In some examples, object analyzer 230 may access shared memory 240. In this example, first device 210 writes to shared memory 240. In this example, second device 220 reads from shared memory 240. Thus, second device 220 can obtain data by reading data from shared memory 240. In this example, first device 210 is communicatively coupled to second device 220 via shared memory 240. In this example, first device 210 is indirectly communicatively coupled to second device 220 via shared memory 240.

[0028] In this specific example, besides being indirectly coupled to the second device 220 via shared memory 240 in a communicative manner, the first device 210 is also directly coupled to the second device 220. In some instances, the second device 220 uses a direct connection with the first device 210 to transmit control data to the first device 210. However, in some instances, the first device 210 is not directly coupled to the second device 220.

[0029] In some instances, the first device 210 provides encoded data to the second device 220. For example, the first device 210 may include encoder functionality, and the second device 220 may include decoder functionality. In other instances, the data provided by the first device 210 to the second device 220 is not in an encoded form.

[0030] refer to Figure 3 An example of system 300 is shown. Figure 3 The instance system 300 described in the above reference includes Figure 1 and 2 The corresponding elements in the described example systems 100 and 200 are the same or similar to several other elements. Such elements are indicated using the same reference numerals as those in 200 and 100, but with corresponding additions.

[0031] In this example, system 300 includes a hierarchical system 300. The hierarchical system 300 can be used to represent signals conforming to a structured hierarchy. Such signals are collectively referred to below as "hierarchical structured signals." In this example, constructing a hierarchical signal involves constructing the signal according to a hierarchical hierarchy of representations (also called "reproductions" or "versions"). In this example, each representation of the signal is associated with a corresponding quality level ("LoQ"). Thus, in this example, the hierarchical system 300 processes data represented according to a hierarchical hierarchy, wherein the hierarchical hierarchy contains multiple different LoQs.

[0032] In some instances, signals are encoded and decoded within the hierarchical system 300. In other instances, signals are not encoded and decoded within the hierarchical system 300.

[0033] For convenience and simplicity, in this specific example, the signal is encoded and decoded within the hierarchical system 300. Therefore, in this example, the hierarchical system 300 includes a hierarchical coding system 300. In this specific example, in addition to constructing the signal hierarchically, the signal is also encoded and decoded within the hierarchical coding system 300. Therefore, the signal encoded in the hierarchical coding system 300 can be considered a "hierarchically coded signal." Thus, in this example, the hierarchical coding system 300 encodes a hierarchically structured signal. Compared to at least some other coding techniques, hierarchical coding produces different levels of compressed images. In this example, the hierarchical layer includes at least two layers (also called "levels"). The hierarchical layer can add detail layers and amplification on top of industry-standard codecs. Examples of such codecs include, but are not limited to, H.264 and High Efficiency Video Decoding (HEVC). Alternatively, the hierarchical layer can be a 'full-stack' layer that does not depend on other standard codecs.

[0034] Readers may refer to international (PCT) patent applications PCT / IB2014 / 060716, PCT / IB2012 / 053660, PCT / IB2012 / 053722, PCT / IB2012 / 053723, PCT / IB2012 / 053724, PCT / IB2012 / 053725, PCT / IB2012 / 056689, PCT / IB2012 / 053726, PCT / GB2018 / 053551, PCT / GB2018 / 053552 and PCT / GB2018 / 053553, which further describe how data is processed in such a hierarchical manner, and are hereby incorporated herein by reference.

[0035] In this example, the first device 310 receives a representation of a signal at a relatively high resolution from a source. In this example, the relatively high resolution corresponds to the resolution of LoQ0. In some examples, the first device 310 receives the representation 351 at LoQ0 directly from the source. In some examples, the first device 310 receives the representation 351 at LoQ0 indirectly from the source. In some examples, the source includes electronic devices that generate and / or record signals using one or more sensors. For example, in the case of video signals, the electronic devices may include a camera. The camera may record a scene at a specific relatively high LoQ (e.g., ultra-high definition (UHD)). The camera may use several sensors (e.g., charge-coupled device (CCD), complementary metal-oxide-semiconductor (CMOS), etc.) to capture information in the scene (e.g., light intensity at a specific location in the scene). The video signal may be further processed and / or stored (via the camera and / or another device) before being received by the first device 310.

[0036] In this specific example, the signal is in the form of a video signal including a sequence of images. However, it should be understood that the signal can be of different types. For example, the signal may include an audio signal, a multi-channel audio signal, a picture, a two-dimensional image, a multi-view video signal, a 3D video signal, a radar / LiDAR and other sparse data signal, a volumetric signal, a volumetric video signal, a medical imaging signal, or a signal with more than four dimensions.

[0037] In this example, the hierarchical system 300 provides representations of video signals at multiple different LoQs. One measure of the LoQ of a video signal representation is its resolution or number of pixels. Higher resolution corresponds to a higher LoQ. Resolution can be spatial and / or temporal. Another measure of the LoQ of a video signal representation is whether the representation is gradient or interlaced, where gradient corresponds to a higher LoQ compared to interlacing. Another measure of LoQ can be the relative quality of the video signal representation. For example, two representations of a video signal may have the same resolution, but one may still have a higher LoQ than the other because it provides more detail (e.g., more edges and / or contours) or better bit depth (e.g., HDR instead of SDR).

[0038] In this example, the first device 310 generates a first set of representations 350 of the signal. In this example, the first set of representations 350 includes at least two representations of the signal. In this specific example, the first set of representations 350 includes 'X+1' representations of the signal. In this example, each of the 'X+1' representations of the signal has an associated LoQ (or "at the associated LoQ"). In this example, one of the 'X+1' representations 351 has LoQ '0' (collectively referred to below as "at LoQ 0"), and another of the 'X+1' representations 352 has LoQ '0'. -1 The same applies below, until 'X+1' represents yet another one of 353 in LoQ. -(X-1) Up to this point, and the 'X+1' represents the last of 354 in LoQ. -X Below. The first group indicates that 350 is available in LoQ. -1 The following represents 352 and LoQ. -(X-1) The following representations 353 may include one or more additional representations. Although in this example, the first set of representations 350 includes at least four representations of the signal, in other examples, the first set of representations 350 may include more or fewer representations of the signal than four.

[0039] The first device 310 can generate the first set of representations 350 in various different ways. In this example, the signal includes a video signal. In this specific example, the representation 351 at LoQ0 corresponds to the original representation of an image included in the video signal. The image can be considered as a time sample representing the video signal. In this example, the first device 310 generates increasingly lower LoQs (i.e., below LoQ) by: 0的 Each of the representations 352, 353, and 354 under LoQ is generated by downsampling (also known as "scaling down") the representation 351 under LoQ0. -1 The lower representation is 352, obtained through downsampling LoQ. -1 The following represents 352 to generate LoQ.-(X-1) The sub-representation 353 (possibly generating one or more intermediate representations via subsampling) is passed through subsampling LoQ. -(X-1) The following represents 353 to generate LoQ. -X The representation below is 354. In this example, undersampling reduces the resolution of the representation. Therefore, in this example, the resolution of the signal representations 351, 352, 353, and 354 decreases as LoQ decreases from '0' to '-X'. As a particular non-limiting example, in a four-layer hierarchy, representation 351 at LoQ0 can correspond to a high-resolution image (e.g., 8K at 120 frames per second), LoQ... -1 The following representation of 352 corresponds to a medium-resolution image, LoQ. -2 The following representation of 353 corresponds to a low-resolution image, and LoQ... -3 The following representation of 354 may correspond to an image with a thumbnail resolution (e.g., below standard definition (SD)).

[0040] Therefore, in this example, each of representations 531, 352, 353, and 354 relates to the same time sample (i.e., image) of the video signal. Thus, this example differs from a set of representations of different time samples of the video signal. For example, a set of representations of different time samples of the video signal includes a first representation of the video signal at a first time, a second representation of the video signal at a second time, and so on. Although in this example, such representations of different time samples of the video signal are associated with corresponding different LoQs, these representations can all be under the same LoQ.

[0041] In this example, the first device 310 will LoQ -X The following representation 354 is transmitted to the second device 320 (e.g., in LoQ). -X (The following indicates that 354 is not encoded), and / or transmit data, so that the second device 320 obtains LoQ. -X The following represents 354 (e.g., LoQ). -X The following represents the encoded version of 354, which can be decoded by the second device 320 to obtain LoQ. -X The following is represented as 354).

[0042] It should be understood that representations 351, 352, 353, and 354 in the first set of representations 350 correspond to the corresponding representations 361, 362, 363, and 364 in the second set of representations 360. When the first device 310 encodes the first set of representations 350 and when the encoding performed by the first device 310 is lossless, representations 351, 352, 353, and 354 in the second set of representations 350 are identical to the corresponding representations 361, 362, 363, and 364 in the first set of representations 360. In other instances where the first device 310 encodes the first set of representations 350, representations 351, 352, 353, and 354 in the first set of representations 350 correspond to the corresponding representations 361, 362, 363, and 364 in the second set of representations 360, but may not be identical to them.

[0043] In this example, LoQ is transmitted as the first device 310. -X The following representation 354 and / or data enables the second device 320 to obtain LoQ -X The result of 354 is shown below, and the second device 320 obtains the LoQ. -X The following represents 364. As noted above, in some instances, LoQ... -X The following represents 364 and LoQ -X The representation below is the same as 354, but it is different in other instances.

[0044] In this example, the second device 320 can 'reverse' the processing performed by the first device 310 to obtain some of all the second set of representations 360. In some instances, the second device 320 samples (or "amplifies") the LoQ. -X The following represents 364 to obtain LoQ. -(X-1) The following represents 363, and so on, up to LoQ. -1 The representation 362 below is sampled up to obtain the representation 361 below LoQ0. Amplification can be achieved using statistical methods. Statistical methods can use, for example, an averaging function. Therefore, the second device 320 can use LoQ... -X The following indicates that 364 obtained the LoQ. -(X-1) The following represents 363, where LoQ -X The following uses of 364 include scaling LoQ. -X The representation below is 364. Although the second device 320 can reverse this process 'completely' to obtain the representation 361 below LoQ0, in some instances, the second device 320 does not obtain the entire second set of representations 360. For example, the second device 320 can amplify up to LoQ0. -1 or below LoQ -1 The LoQ, but not higher than them.

[0045] Therefore, in this example, the second device 320 obtains a first representation of the signal at the first LoQ, i.e., LoQ. -X The following represents 364. The first representation is LoQ. -X The representation 364 below is a portion of a set of representations of the signal, namely the second set of representations 360. In this example, the set of representations, namely the second set of representations 360, includes a second representation of the signal at a higher second LoQ. The second representation may include LoQ. -(X-1) The following represents 363 and LoQ. -1 The representation below LoQ is 362, or the representation below LoQ0 is 361. In this example, the set of representations, namely the second set of representations 360, includes a third representation of the signal at a third LoQ higher than the second LoQ. The third representation may include LoQ. -1 The representation below LoQ is 362, or the representation below LoQ0 is 361. Therefore, in this example, the second set of representations 360 includes multiple representations of the signal at corresponding LoQs, where each LoQ is equal to or higher than LoQ. -X .

[0046] In some instances, the first device 310 also transmits reconstructed data (also referred to as "residual data" or "residue") to the second device 320. The reconstructed data allows the second device 320 to compensate for inaccuracies in the upsampling process performed by the second device 320, thereby obtaining a more accurate reconstruction of the first set of representations 350. Specifically, the second device 320 may amplify a LoQ (e.g., LoQ...) -X The representation below is used to generate a higher-level LoQ (e.g., LoQ). -(X-1) The approximation (or "prediction") of the representation is given by the first device. The reconstructed data can be used to adjust the approximation to address the aforementioned inaccuracies. The first device 310 can deliver a complete set of reconstructed data, enabling lossless reconstruction. The first device 310 can deliver quantized reconstructed data, which enables visually lossless or lossy reconstruction. The reader may refer to PCT / IB2014 / 060716, which describes the reconstructed data in detail.

[0047] In this example, object analyzer 330 includes at least one object analysis element 370. In this example, object analysis element 370 is configured to perform object analysis within a second set of representations 360. In this example, second device 320 includes 'X+1' object analysis elements 371, 372, 373, and 374. In this example, the number 'X+1' of object analysis elements 371, 372, 373, and 374 is the same as the number of LoQs in the hierarchical hierarchy upon which the signal is encoded in the hierarchical system 300. However, in other examples, the number of object analysis elements differs from the number of LoQs.

[0048] In this example, one object analysis element 371 is at LoQ0, and another object analysis element 372 is at LoQ0. -1 Next, another object analysis element 373 in LoQ -(X-1) Below, and the last object analysis element 374 in LoQ -X Below. In this example, object analysis elements 371, 372, 373, and 374 execute LoQ0 and LoQ respectively. -1 LoQ -(X-1) and LoQ -X The following object analysis. For example, object analysis components 371, 372, 373, and 374 may have been trained to perform LoQ0, LoQ... -1 LoQ -(X-1) and LoQ -X One or more types of object analysis are described below in more detail. For example, object analysis elements 371, 372, 373, and 374 may have been trained using training images that have LoQ0, LoQ, and LoQ respectively. -1 LoQ -(X-1) and LoQ -X The associated resolution. In this example, object analysis elements 371, 372, 373, and 374 are optimized to perform LoQ0 and LoQ respectively. -1 LoQ -(X-1) and LoQ -X Object analysis is deployed. Therefore, in this example, the first object analysis element, namely LoQ... -X The object analysis element 374 below, and the first LoQ, i.e., LoQ -X Related, the second object analysis element, namely LoQ -(X-1) The object analysis element 373 below, and the higher second LoQ, i.e., LoQ -(X-1) Associated. In some instances, at least one LoQ does not have an associated object analysis element. For example, performing object analysis under said at least one LoQ may not be effective. In some instances, each LoQ has at least one associated object analysis element.

[0049] In this example, LoQ -X The object analysis element 374 is communicatively coupled to LoQ. -(X-1) The object analysis element 373 is described below. For example, it is derived from LoQ. -X The data output from the object analysis element 374 can be provided to LoQ. -(X-1) The object analysis element 373 below. Compared to not being coupled to LoQ in a communicative manner. -X LoQ of the object analysis element 374 -(X-1)The object analysis element 373 below, in turn, can enhance LoQ. -(X-1) The performance of the object analysis element 373 is analyzed below. For example, by LoQ -X The results of object analysis performed by the object analysis element 374 and / or the results of LoQ -X The object analysis data performed by the object analysis element 374 can be provided to LoQ. -(X-1) The object analysis element 373 is located below. LoQ -(X-1) The object analysis element 373 below can, in turn, perform LoQ. -(X-1) This type of data is used during object analysis. As a result, data from LoQ is used... -X A supplement or alternative to the data of the object analysis element 374, LoQ -(X-1) The object analysis element 373 below can be used with other LoQ components. -X Related data. For example, LoQ -(X-1) The object analysis element 373 below can provide LoQ -X The following represents 364 and / or from LoQ -X The following represents the data obtained from LoQ. Compared to not using LoQ... -X Related data, using this type of data can improve performance in LoQ -(X-1) The accuracy of the object analysis performed below, because in addition to the LoQ relative to the object analysis it performs, -(X-1) The following indicates LoQ in addition to 363. -(X-1) The object analysis element 373 below also has additional data. In this example, LoQ -0 The object analysis element 371 is communicatively coupled to LoQ. -1 The object analysis element 372 below.

[0050] In some instances, object analysis element 370 performs object analysis relative to one or more representations of a signal at a given time sample, and the result of the object analysis performed relative to the given time sample is used to perform object analysis relative to another time sample of the signal. This other time sample of the signal can be a subsequent time sample of the signal. For example, object analysis performed relative to one image in an image sequence can be used to influence object analysis performed relative to one or more subsequent images in the image sequence. This can enhance the object analysis performed relative to one or more subsequent images.

[0051] Therefore, object analysis can be performed within one or more LoQs in the second set of representations 360 within the hierarchical system 300. Thus, in such instances, object analysis is performed within the hierarchical system, specifically within the instance hierarchical system 300.

[0052] Object analysis elements can take different forms. In some instances, object analysis elements include convolutional neural networks (CNNs). In some instances, object analysis elements include multiple CNNs. Thus, object analysis can be performed using one or more CNNs. CNNs can be trained to perform object analysis in various different ways relative to the representation of the signal. For example, CNNs can be trained to detect and locate one or more objects, and / or identify one or more objects within the representation of the signal. Although in some instances object analysis elements include one or more CNNs, in other instances, hierarchical applications of long short-term memory (LSTM) or dense neural networks (DNNs) can be used. In some instances, object analysis elements may not include artificial neural networks (ANNs). For example, discrete optimizers can be used.

[0053] In some instances, performing object analysis includes detecting objects. Thus, object analysis element 370 can be configured to perform object detection within one or more of representations 361, 362, 363, and 364 included in the second set of representations 360. In such instances, object detection is performed within a hierarchical system 300. Object detection involves finding one or more instances of one or more objects of one or more specific categories and locating said one or more objects within said representation. Therefore, object detection may involve detecting all objects belonging to a specific category for which object analysis element 370 has been trained, and locating them within the representation. For example, object analysis element 370 may be trained to detect faces and animals. If such object analysis element 370 detects one or more such objects, then the specific location of each such detected object is returned, for example, via bounding boxes. For example, bounding boxes may be provided relative to each face and each animal detected in the representation. The result of such object detection may be a confidence level that one or more objects have been detected and located.

[0054] In some instances, performing object analysis includes identifying objects. Therefore, the object analysis element 370 can be configured to perform object identification within one or more of the representations 361, 362, 363, and 364 included in the second set of representations 360. In such instances, object identification is performed within a hierarchical system 300. Object identification involves recognizing detected objects. For example, object identification may involve determining the category label to which a detected animal belongs, such as 'dog,' 'cat,' 'Persian cat,' etc. In the case of faces, object identification may involve identifying the identity of a specific person whose face has been detected. For example, the result of such object identification could be a given confidence level that the detected animal is a cat. For example, the result of such object identification could correspond to an 80% confidence level that a cat has been identified and a 20% confidence level that a cat has not been identified.

[0055] Therefore, one or more different types of object analysis include, but are not limited to, object detection and object recognition, which can be performed within the second set of representations 360.

[0056] In some instances, object analysis of a given type (e.g., object detection) is performed only within a single LoQ within a second set of representations (360). For example, object detection may be performed only within a LoQ. -X The lower LoQ representation executes within 364, but not within any of the higher LoQ representations 361, 362, or 363 in the second group of representations 360. This can be achieved, for example, in LoQ... -X The object analysis element 374 is performed when the object has been detected and located. In this case, at one or more higher LoQs, such as at LoQ... -(X-1) Performing further object detection may not be an efficient and effective use of the resources of the object analyzer 330, since the object has already been detected and located.

[0057] Although in this example, object analysis begins at the lowest LoQ, i.e., LoQ -X However, in other instances, object analysis can begin at a higher LoQ. Starting with the lowest LoQ saves the extra processing time and / or resources that contribute to successful object analysis at the lowest LoQ. In some instances, object analysis at the lowest LoQ may not significantly contribute to successful object analysis. In such instances, starting object analysis at a higher LoQ may be more efficient than starting at the lowest LoQ.

[0058] In some instances, object analysis proceeds along an ascending hierarchy, where object analysis is performed at progressively higher LoQs. In other instances, object analysis may involve proceeding along a descending hierarchy. For example, although the hierarchy may initially be ascending, object analysis may be performed at one or more LoQs below the current LoQ. Object analysis may or may not have been performed at said one or more lower LoQs.

[0059] In some instances where multiple different types of object analysis are performed, some or all of these analyses begin with the same LoQ. In other instances where multiple different types of object analysis are performed, some or all of these analyses begin with different LoQs. This may be particularly effective when a given type of object analysis is more efficient at a higher or lower LoQ, but it is not the only effective approach.

[0060] In some instances, object analysis of a given type (e.g., object detection) is performed within multiple LoQs within a second set of representations 360. For example, object detection can be performed relative to a LoQ. -XThe following indicates that 364 will execute, but LoQ... -X The object analysis element 374 may not have detected an object. In this case, object detection can be performed by LoQ. -(X-1) The object analysis element 373 under LoQ -(X-1) Execute below. If LoQ -(X-1) If object analysis element 373 does not detect an object, then object detection can be performed again at one or more higher LoQs. Although in this instance, object detection is performed at one LoQ (i.e., LoQ...). -X If no object is detected, the object detection is performed at the next higher level, LoQ (i.e., LoQ). -(X-1) Object detection and object identification can be performed at a low LoQ, but in other instances, one or more higher LoQs can be bypassed (or “skipped”). This can be particularly effective, but not the only effective, situation when object analysis of a given type is not efficient at a relatively low LoQ but may be efficient at one or more higher LoQs. Object detection and object identification can operate efficiently simultaneously at multiple LoQs. If multiple objects exist, some objects can be detected with sufficient accuracy at a low LoQ, while others can be detected at higher positions in the hierarchy. The same applies to object identification. Some features can be identified at a low LoQ (e.g., humans), some at a higher LoQ (e.g., men wearing glasses), and some at a high LoQ (e.g., facial recognition of a specific individual). Different identification problems can be solved at different LoQs. For example, as the hierarchy increases, the identification problem may be refined to obtain a more detailed description of the object. For example, at a low LoQ (e.g., LoQ 1), the identification problem may be refined to obtain a more detailed description of the object. -5 Under these conditions, the question might be whether anyone has a higher LoQ (e.g., LoQ 0). -4 Under this condition, the problem might be the person's gender, at a higher LoQ (e.g., LoQ). -3 Under these conditions, identity, emotions, etc., are identified again.

[0061] Specific, non-limiting examples will now be provided, in which object analysis element 374 is in LoQ -X The following indicates that object identification is performed within 364 seconds. In this specific instance, a hypothesis is tested. For example, the hypothesis could be a cat in LoQ. -X The following is represented in 364. Object analysis element 374 may have been trained to identify LoQ. -X The cat within the representation (e.g., an image) below. LoQ -X The object analysis element 374 can be based on the analysis LoQ -X The following represents 364 to determine the result of the hypothesis. For example, LoQ -X The object analysis element 374 below can determine that the assumption is correct (i.e., LoQ). -XThe confidence level for the following (representing the existence of cats in 364) is, for example, 20%, with the assumption being incorrect (i.e., LoQ). -X The confidence level for (indicating that cats do not exist in 364) is, for example, 80%. This result indicates that LoQ -X The representation 364 below the LoQ is extremely unlikely to contain a cat. However, this does not mean that there are actually no cats in any of the representations 361, 362, and 363 below the higher LoQ in the second set of representations 360. For example, the presence of an animal may be obvious at a lower LoQ, but an animal specifically a cat may only become obvious at a higher LoQ. In some instances, the confidence level of the hypothesis being correct is compared to a threshold confidence level. In this particular instance, the threshold confidence level is assumed to be 50%. In this instance, the determined 20% confidence level of the hypothesis being correct is lower than the 50% threshold confidence level. Therefore, in this instance, the confidence level associated with object identification performed within representation 364 below LoQ-X does not meet the threshold confidence level. Such a threshold confidence level may be called an "object identification confidence threshold level" because it can be used as a threshold for the confidence level associated with object identification.

[0062] In response to determining that the confidence level is below a threshold confidence level, one or more predetermined actions may be taken.

[0063] One example of such a pre-defined action is to stop (also known as "abort") the execution of the second set of object analysis representing 360 degrees. For example, LoQ... -X A low confidence level indicates that object analysis is unlikely to succeed within the second set of representations (360 degrees). This can save the additional resource usage associated with continuing object analysis where the probability of success is low. In some instances, a specific 'stop object analysis' threshold can be configured. This stop object analysis threshold can correspond to an extremely low probability of object analysis success.

[0064] Another example of such pre-defined actions is continuing the analysis of that specific type of object within the second set of representations (360°). For example, with LoQ... -X A low confidence level associated with object identification performed at a lower LoQ may not mean that object identification will not succeed within 360 degrees of the second representation. For example, object identification may be more effective at a higher LoQ. Therefore, a confidence level for object identification performed at a lower LoQ does not preclude success at a higher LoQ. For example, if at a lower LoQ... -X If the confidence level of object identification performed is lower than the threshold confidence level, then it can be performed in LoQ. -(X-1) The following indicates object identification within LoQ 363. -(X-1) The confidence level of object identification performed below can be compared with that in LoQ. -XThe same threshold confidence level is used for comparison. Alternatively or additionally, in LoQ... -(X-1) The confidence level of the object identification performed below can be compared with another threshold confidence level. In some instances, the other threshold confidence level is compared with LoQ. -(X-1) Related to and not related to LoQ -X The threshold confidence level is associated with the following. In some instances, the other threshold confidence level is associated with LoQ. -X This is associated with a threshold confidence level. For example, the other threshold confidence level could be LoQ. -X Threshold confidence level and LoQ -(X-1) A function of the threshold confidence level. An example of such a function is the average. Object identification can be performed above LoQ. -(X-1) Execute under one or more LoQs.

[0065] In some instances, as a supplement or alternative to comparing confidence levels with threshold confidence levels, the number of rising LoQs can be compared with the threshold number of rising LoQs. For example, when the threshold number of rising LoQs is three, if object analysis of a given type has already been performed at three LoQs without success, then the object analysis of that given type can be abandoned. In such instances, resource usage associated with object analyzer 330 and / or second device 320 can be saved compared to continuing object analysis of the given type, where the probability of success is low.

[0066] In response to determining that the confidence level is higher than the threshold confidence level, one or more predetermined actions may be taken.

[0067] One instance of this pre-defined action ends within 360 degrees of the second group of representations of object analysis. In some instances, ending object analysis includes ending all types of object analysis. In some instances, ending object analysis includes ending one or more of a given type of object analysis and starting or continuing one or more other types of object analysis. For example, in LoQ... -X The confidence level of the object identification performed below indicates that the object has been identified with sufficient confidence, thus eliminating the need to perform further object identification at one or more higher LoQs. This can save resources associated with the second device 320 and / or the object analyzer 330.

[0068] Another example of such pre-defined actions is continuing object analysis within the second set of representations (360 degrees). In some instances, continuing object analysis includes continuing object analysis of all types. In some instances, continuing object analysis includes continuing object analysis of one or more given types. For example, in LoQ... -XThe confidence level of object identification performed at a given LoQ indicates that objects have been identified with a relatively high degree of confidence. However, object identification can be performed at one or more higher LoQs to refine the LoQ level. -X The confidence level of the object identification performed at a higher LoQ. Performing object identification at one or more higher LoQs can change the confidence level, for example, by increasing and / or decreasing it. Using the non-limiting example of identifying a cat described above, the established assumption is correct (i.e., the cat is at a higher LoQ). -X The confidence level of 20% (as shown in LoQ) can be lower than the threshold confidence level of 50% for assuming the hypothesis is correct. However, based on LoQ... -(X-1) The object identification performed below, assuming correctness, can increase the confidence level to 55%, higher than the 50% confidence level threshold for correctness. The confidence level can be increased in this way, for example, compared to LoQ. -X The object analysis element 374 below, the existence of the cat for LoQ -(X-1) The object analysis element 373 below explains this much more clearly. This can be LoQ. -(X-1) The result of increased resolution.

[0069] Therefore, object analysis can be performed more efficiently with respect to hierarchical signals than with other types of signals, based on the examples described herein. In this example, object analysis element 374 can be performed at the lowest LoQ, i.e., LoQ -X Analyze the execution object. LoQ -X Representation 364 below is a lower LoQ representation of the signal, different from representation 361 below LoQ0. As noted above, representation 361 below LoQ0 corresponds to the original source representation of the signal, at least in terms of resolution. Therefore, Figure 3 The instance system 300 depicted differs from systems where object analysis is performed within a representation of a signal, wherein the representation is not part of a set of representations of the signal (e.g., at different image resolutions). In contrast, Figure 3 The example system 300 depicted includes a hierarchical system 300, wherein signals are constructed according to a hierarchical structure comprising multiple different representations of the signals at multiple different corresponding LoQs. Performing object analysis within a representation at a relatively low LoQ in the hierarchical structure allows object analysis to be faster than analysis performed within the representation of the signal in a subset of representations that are not signals. One reason for this is that the second device 320 obtains the LoQ. -X The amount of time required for representation at LoQ0 (also known as "latency") is less than the time required for the second device 320 to obtain the representation at LoQ0. This is because the second device 320 is designed for LoQ0. -X The data volume obtained under LoQ is less than that under LoQ0. If it is possible to obtain data under LoQ... -XIf object analysis is successfully performed, the time required to perform such object analysis may be less than the time required to perform object analysis under LoQ0. One factor to consider is that the first device 310 generates the LoQ. -X The following represents the amount of time involved in the processing performed by the first device 310, which is LoQ. -X The following represents 354, LoQ -X The following indicates that 354 is transmitted to the second device 320 and the second device 320 successfully performs object analysis (which may involve the second device 320's LoQ). -X The time involved in upsampling all or part of the representation 364 under LoQ0 is less than the time involved in the first device 310 transmitting the representation 351 under LoQ0 to the second device 320, and the second device 320 successfully performs object analysis under LoQ0, thus saving processing time. This time-saving method may be particularly effective in real-time systems, but it is not the only effective one; in real-time systems, a reduction in processing time can improve performance. For example, a reduction in processing time can improve the performance of computer vision systems. In some instances, upsampling LoQ0 is more efficient than upsampling all or part of the representation 351 under LoQ0 to the second device 320. -X The reduced amount of data involved in transmitting data from the first device 310 to the second device 320 (represented by the instruction 354) represents a significant saving in processing time, even if the processing performed by the first device 310 is time-intensive. For example, the first device 300 may generate a LoQ. -X The following is represented as 354 and stored in shared memory (240; Figure 2 At a later time, the second device 320 may access the shared memory (240; Figure 2 Searching for LoQ -X The following represents 354, in LoQ -X The following indicates that object analysis is performed within 364 seconds. Given the relatively small amount of data to be retrieved, LoQ is retrieved in the second device 320. -XIn the case of representation 354 at LoQ 0, the retrieval time may be shorter than in the case of the second device 320 retrieving representation 351 at LoQ 0. In some instances, performing object analysis at a lower LoQ may actually be more efficient than performing object analysis at a higher LoQ. For example, denoising may occur at a lower LoQ relative to a higher LoQ. Object analysis may be more efficient in the denoised representation of a signal, for example, even at lower resolutions. Performing object analysis at multiple different LoQs can enhance object analysis compared to performing it at a single LoQ. For example, object analysis elements (e.g., including CNNs) can identify different features at different LoQs. For example, different features can be identified at different resolutions. For instance, a representation of a signal at a relatively low LoQ can provide a complete picture of the scene represented by the signal. Object detection can be performed efficiently relative to such a representation of the signal. For example, humans can be detected, located, and identified at a low LoQ, which may trigger an object avoidance procedure. In this instance, additional details, such as whether the human is male or female, or whether they are wearing sunglasses, may not be required; these details can be identified using further object analysis. For example, this further detail may only become apparent at higher LoQs.

[0070] Object detection and localization can help define one or more RoIs. This allows further object analysis, such as object identification, to be performed within one or more RoIs, rather than within the entire representation. Therefore, for further object analysis, portions of the representation outside the RoI can actually be discarded. For example, if an object has already been detected and located at a relatively low LoQ, further object analysis at a relatively low and / or intermediate LoQ and / or relatively high LoQ constrained to the RoI can be used to adjust the object avoidance procedure. For instance, if the detected object is determined to pose no collision risk, for example, by identifying the object's properties based on object identification, then the object avoidance procedure can be aborted.

[0071] refer to Figure 4 This shows an instance of a portion of instance system 400. Figure 4 The instance system 400 described in the above reference includes Figure 1 , 2 Several elements are the same as or similar to the corresponding elements in the example systems 100, 200, and 300 described in 3. Such elements are indicated by the same but correspondingly increased reference numerals for 300, 200, and 100.

[0072] In this specific instance, the second group represents 460 including LoQ. -X The single representation below is 464. In this example, LoQ... -XThe object analysis element 474 is already in LoQ -X The following indicates that object analysis is performed within 464. In this example, it is performed by LoQ. -X The result of the object analysis performed by the object analysis element 474 is that the object has been detected and located, as shown by LoQ. -X The following indicates RoI 490 for 464. In this example, RoI corresponds to LoQ. -X The sub-region below represents 464. Using the above non-limiting example, object recognition for cats can be performed in LoQ. -X The following indicates execution in RoI 464, and LoQ is detected within RoI 490. -X The following represents animals in RoI 464. The detected animal could be a cat. In this example, RoI 490 represents LoQ. -X The value below represents the part of interest in 464. In this example, LoQ -X The part of RoI 490 in the following representation 464 is of interest because it contains an animal, which may be a cat.

[0073] In this example, LoQ -X The object analysis element 474 performs object identification within RoI 490. In this example, it is assumed that LoQ... -X The object analysis element 474 below did not identify the animal in RoI 490 as a cat.

[0074] In this example, decoder 420 obtains a partial representation of the signal 480. This is consistent with the reference above. Figure 3 The described instance system 300 differs from this one, in which all representations 361, 362, 363, and 364 obtained by the decoder 320 are acquired in their entirety. In this instance, the second device 420 only amplifies the LoQ. -X The following represents the portion of RoI 490 in 464 to obtain the LoQ. -(X-1) The lower part represents 483. In this specific instance, LoQ -(X-1) The lower part indicates that 483 contains detected animals. In this example, LoQ -(X-1) The representation below 483 is a partial representation because LoQ -(X-1) The complete representation below (363; Figure 3 (This information has not yet been obtained.) Therefore, in this example, the second device 420 uses LoQ. -X The following represents the part representing 464 to obtain the LoQ. -(X-1) The lower part represents 483. In this example, the second device 420 does not amplify the LoQ. -XThis applies to the portion of representation 464 under LoQ-X that is outside RoI 490. This may be the case even if object analysis has already been performed on the portion of representation 464 under LoQ-X that is outside RoI 490. In this example, LoQ... -(X-1) The object analysis element 473 under LoQ -(X-1) The following section indicates that object analysis is performed within 483. Therefore, in this example, object identification occurs within LoQ. -(X-1) The part below indicates that the value is within 483 and not in LoQ. -(X-1) The complete representation below (363; Figure 3 Executed within ) . In this particular instance, such object identification involves recognizing whether the detected animal is a cat. In this instance, the second device 420 amplifies the LoQ -(X-1) The lower part representation 483 is used to obtain the lower part representation 481 below LoQ0. For example, object recognition may not succeed at levels prior to LoQ0. Such amplification may involve obtaining one or more intermediate part representations of the signal. Such amplification may involve recognizing LoQ... -(X-1) The lower part represents another RoI in 483, and the LoQ is magnified. -(X-1) The lower portion represents the portion of 483 corresponding to the RoI. In this example, the object analysis element 471 at LoQ0 performs object analysis within the portion of representation 481 at LoQ0. Although in this example, the second device 420 partially processes (e.g., partially decodes) to LoQ0, in other examples, the second device 420 stops at a lower LoQ. This can be considered to correspond to “scaling” or “clipping” within one or more representations of the signal relative to one or more RoIs.

[0075] In some instances, the second device 420 receives control data for controlling, for example, whether full processing or partial processing (e.g., decoding) should be performed at a given LoQ. In some instances, the second device 420 receives control data directly from the object analyzer 430. In some instances, the second device 420 receives control data from an intermediate element between the second device 420 and the object analyzer 430. An example of such an intermediate element is a dedicated RoI identification element capable of identifying one or more RoIs.

[0076] In this specific instance, only one object has been detected and located, and only one RoI 490 has been identified. In some instances, more than one object has been detected and located. In such instances, multiple RoI 490s may be used.

[0077] Therefore, in this example, object analysis can be performed by combining partial representations with the complete representation. This partial processing can reduce processing time compared to full processing. Furthermore, such partial processing may involve the second device 420 acquiring and processing less data than full processing to enable object analysis to be performed. The hierarchical structure of the signal according to the example described herein is particularly well-suited for partial processing. In fact, such partial processing allows navigation to the portion of the representation of interest.

[0078] Therefore, based on the examples described herein, object analysis can use signals in the first LoQ (e.g., LoQ). -X The representation of the signal at the first LoQ is at least partially executed, and the representation of the signal at the first LoQ is performed using the signal at a higher second LoQ (e.g., LoQ). -(X-1) The representation is generated at (and higher) levels, where performing object analysis includes performing object detection and / or object identification. Performing such object analysis using at least a portion of the signal's representation at a first LoQ may include performing object analysis within the signal's representation at a first LoQ and / or obtaining one or more complete and / or partial representations of the signal at one or more higher LoQs, and performing object analysis within the one or more complete and / or partial representations of the signal at the one or more higher LoQs.

[0079] Furthermore, based on the examples described herein, object analysis can be performed within images in a multi-resolution image format, where multiple versions of the image at different corresponding image resolutions are available. The multi-resolution image format can correspond to a format where the image has a hierarchical structure as described herein.

[0080] Furthermore, according to the examples described herein, the representation can be partially processed (e.g., partially decoded, partially reconstructed, partially amplified) in response to object analysis performed within the representation of the signal to identify candidate objects in the RoI within the representation of the signal, wherein the partial processing is performed relative to the RoI.

[0081] refer to Figure 5 A schematic block diagram of an example of device 500 is shown.

[0082] Examples of device 500 include, but are not limited to, mobile computers, personal computer systems, wireless devices, base stations, telephone devices, desktop computers, laptop computers, notebook computers, netbook computers, mainframe computer systems, handheld computers, workstations, network computers, application servers, storage devices, consumer electronic devices such as cameras, portable camcorders, mobile devices, video game consoles, handheld video game devices, peripheral devices such as switches, modems, routers, vehicles, etc., or generally any type of computing or electronic device.

[0083] In this example, device 500 includes one or more processors 501 configured to process information and / or instructions. The one or more processors 501 may include a central processing unit (CPU). The one or more processors 501 are coupled to a bus 502. Operations performed by the one or more processors 501 may be implemented via hardware and / or software. The one or more processors 501 may include multiple processors located in the same location or multiple processors located in different locations.

[0084] In this example, device 500 includes computer-available volatile memory 503 configured to store information and / or instructions for the one or more processors 501. Computer-available volatile memory 503 is coupled to bus 502. Computer-available volatile memory 503 may include random access memory (RAM).

[0085] In this example, device 500 includes a computer-available non-volatile memory 504 configured to store information and / or instructions for the one or more processors 501. The computer-available non-volatile memory 504 is coupled to a bus 502. The computer-available non-volatile memory 504 may include read-only memory (ROM).

[0086] In this example, device 500 includes one or more data storage units 505 configured to store information and / or instructions. The one or more data storage units 505 are coupled to bus 502. The one or more data storage units 505 may, for example, include a magnetic disk or optical disk and a disk drive or solid-state drive (SSD).

[0087] In this example, device 500 includes one or more input / output (I / O) devices 506 configured to transmit information to and / or from the one or more processors 501. The one or more I / O devices 506 are coupled to a bus 502. The one or more I / O devices 506 may include at least one network interface. The at least one network interface enables device 500 to communicate via one or more data communication networks. Examples of data communication networks include, but are not limited to, the Internet and local area networks (LANs). The one or more I / O devices 506 enable users to provide input to device 500 via one or more input devices (not shown). The one or more I / O devices 506 enable information to be provided to users via one or more output devices (not shown).

[0088] Various other entities of device 500 are depicted. For example, operating system 507, data processing module 508, one or more other modules 509, and data 510 (if present) are shown residing in one or a combination of computer-usable volatile memory 503, computer-usable non-volatile memory 504, and the one or more data storage units 505. Signal processing module 508 may be implemented by computer program code stored in a memory location within computer-usable non-volatile memory 504, computer-readable storage media within the one or more data storage units 505, and / or other tangible computer-readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, optical media (e.g., CD-ROM, DVD-ROM, or Blu-ray), flash memory cards, floppy disks, or hard disks, or any other media capable of storing computer-readable instructions such as firmware or microcode in at least one ROM or RAM or programmable ROM (PROM) chip or application-specific integrated circuit (ASIC).

[0089] Therefore, device 500 may include a data processing module 508 executable by the one or more processors 501. Data processing module 508 may be configured to contain instructions that implement at least some of the operations described herein. During operation, the one or more processors 501 initiate, run, execute, interpret, or otherwise perform the instructions in data processing module 508.

[0090] While at least some aspects of the examples described herein with reference to the accompanying drawings include computer processes executed in a processing system or processor, the examples described herein also extend to computer programs, such as computer programs on or within a carrier suitable for putting the examples into practice. A carrier can be any entity or device capable of carrying a program.

[0091] It should be understood that device 500 may include, compared to Figure 5 The more, fewer, and / or different components described in the text.

[0092] Device 500 can be located in a single location or distributed across multiple locations. These locations can be local or remote.

[0093] The techniques described herein can be implemented in software or hardware, or a combination of both. They may include configuring a device to implement and / or support any or all of the techniques described herein.

[0094] The above embodiments should be understood as illustrative examples. Other embodiments are contemplated.

[0095] It should be understood that any feature described with respect to any embodiment may be used alone or in combination with other described features, and may also be used in combination with one or more features of any other embodiment, or in any combination with any other embodiment. Furthermore, equivalents and modifications not described above may be employed without departing from the scope of the invention as defined in the appended claims.

Claims

1. A method for performing object detection within a set of representations of a hierarchical structured signal, wherein, The hierarchical structured signal is constructed according to a hierarchical structure of representations of the original signal, and each representation is associated with a corresponding quality level. The set of representations includes at least a first representation of the original signal at a first quality level and a second representation of the original signal at a second quality level, where the second quality level is higher than the first quality level. The signal includes a video signal, wherein the second representation is based on reconstructed data used to adjust the first representation. The method includes: Object detection is performed in one or more representations to find one or more instances of one or more objects of one or more specific categories and to locate the one or more objects within the one or more representations. The step of performing object detection in one or more representations includes: Object detection is performed in the first representation of the original signal to obtain data associated with the object detection in the first representation; and Object detection is performed in the second representation of the original signal using the data associated with the object detection in the first representation.

2. The method of claim 1, wherein the object detection is performed using at least one convolutional neural network (CNN), wherein the object detection is performed using a first CNN associated with the first quality level and a second CNN associated with the second quality level.

3. The method of claim 2, wherein the data output by the first CNN is provided to the second CNN.

4. The method according to any one of claims 1 to 3, comprising identifying one or more internal execution objects in the set of representations.

5. The method according to any one of claims 1 to 3, comprising obtaining the first representation by decoding one layer of the hierarchical structure signal, wherein, The hierarchical structure signal is encoded into a hierarchical coded signal received from the encoder.

6. The method according to any one of claims 1 to 3, comprising obtaining at least a portion of the second representation using at least a portion of the first representation.

7. The method of claim 6, wherein the at least portion of the second representation is obtained in response to determining that the confidence level associated with object identification performed within the first representation does not meet the object identification threshold confidence level.

8. The method of claim 6, further comprising obtaining only a portion of the second representation using only a portion of the first representation.

9. The method of claim 8, wherein object detection and / or object identification are performed within the portion of the second representation.

10. The method of claim 8, wherein the portion of the first representation corresponds to the region of interest within the first representation.

11. The method of claim 1, wherein the first and second represent samples of the same time, each having the same time of the video signal.

12. The method according to any one of claims 1 to 3, wherein the quality level corresponds to the image resolution.

13. An apparatus configured to perform the method according to any one of claims 1 to 12.

14. A computer-readable storage medium configured to perform the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Signal processing and inheritance in a tiered signal quality hierarchy

    CN103918261A

  • Low- and high-fidelity classifiers applied to road-scene images

    US20170206434A1