Vector database system based on query distribution perception

By introducing distribution-aware and dynamic index optimization mechanisms into vector databases, the performance degradation caused by inconsistent distribution of user queries and vectors in databases is solved, and stable and efficient retrieval of vector database query quality is achieved.

CN120067147AActive Publication Date: 2025-05-30HANGZHOU DIANZI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510126887.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

When existing vector databases deal with the situation where user queries are inconsistent with vector distribution in the database, query performance drops sharply, resulting in high latency and low recall.

Method used

A vector database system based on query distribution perception is designed. Through the index construction module, vector search module, interval measurement module, logging module, distribution perception module, distribution detection module and distribution mapping module, the distribution difference between the query vector and the vector in the database is monitored and adjusted in real time, and the index structure is dynamically optimized to improve the retrieval efficiency.

Benefits of technology

It effectively solves the problems of high latency and low recall caused by user queries of different distributions, ensures the stability and reliability of the query quality of vector databases, and significantly improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067147A_ABST
    Figure CN120067147A_ABST
Patent Text Reader

Abstract

The invention relates to the field of vector databases, and provides a vector database system based on query distribution perception aiming at the problems of high delay and low recall rate caused by different distributed user queries. Comprising an index construction module, a vector retrieval module, an interval measurement module, a log recording module, a distribution perception module, a distribution detection module, a distribution mapping module, a storage module, an interface agent module and a graphical management interface. According to the method, a set of efficient solution is provided for storage and retrieval of vector data, meanwhile, the distribution condition of user query is monitored in real time, vector distribution is effectively measured, a vector index structure is optimized through user historical query, and stability and reliability of the query quality of a vector database are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vector databases, and particularly to a vector database system based on query distribution awareness. Background Art

[0002] With the booming development of artificial intelligence, machine learning, and deep learning, the importance of vector data has become increasingly prominent in numerous fields. However, traditional database technologies have obvious deficiencies in storing and retrieving vector data, so the emerging technology of vector databases has emerged as the times require.

[0003] In the actual application of vector databases, it is extremely common for there to be differences between user queries and the vector distribution in the database. For example, the query vector and the database vector have inconsistent modalities, and the user query distribution changes over time. However, currently, vector similarity retrieval in vector databases is usually based on the assumption that the query vector and the vectors in the database follow the same distribution. Once the distributions of the query vector and the database vectors are different, the query performance will drop sharply, and the performance gap may reach an order of magnitude or more.

[0004] Therefore, the efficient storage and accurate retrieval of vector data are crucial, and the diversity of user query distributions will greatly affect the query quality of vector databases. We must take targeted measures to provide an efficient solution for the storage and retrieval of vector data to meet the query requirements of different distributions. This can not only ensure the stability and reliability of query quality but also significantly improve the user experience, so it has extremely high urgency and importance. Summary of the Invention

[0005] Aiming at the shortcomings and deficiencies existing in the existing vector database technology, the present invention proposes a vector database system based on query distribution awareness. The system consists of an index construction module, a vector retrieval module, an interval measurement module, a log recording module, a distribution awareness module, a distribution detection module, a distribution mapping module, a storage module, an interface proxy module, and a graphical management interface, providing an efficient solution for the storage and retrieval of vector data. At the same time, it monitors the distribution of user queries in real time, visually measures the distribution differences between query vectors and vectors in the database, and effectively solves the problems of high latency and low recall rate caused by user queries with different distributions on the basis of ensuring the retrieval quality of user queries with the same distribution, ensuring the stability and reliability of the query quality of the vector database.

[0006] The present invention proposes a vector database system based on query distribution awareness, including an index construction module, a vector retrieval module, an interval measurement module, a log recording module, a distribution awareness module, a distribution detection module, and a distribution mapping module;

[0007] The index construction module is used to: construct an index for the retrieval target set V = {v 1 , v 2 … v n}, record the index name, where n is the number of vectors in the retrieval target set V;

[0008] The vector retrieval module is used to: receive the query data set Q = {q 1 , q 2 … q m} and the query index name, where m is the number of vectors in the query data set Q; for each element in the query data set Q, query the index according to the query index name to obtain the top-k approximate nearest neighbor retrieval results;

[0009] The interval measurement module is used to: traverse each vector v i in the retrieval target set V, calculate the distance d i between the vector v Mi and the retrieval target set V, form a distance set U(d M ), after the traversal is completed, record the distribution interval of the retrieval target set V, indicating the distribution deviation degree of the retrieval target set V;

[0010] The log recording module is used to: record the query log, and the query log includes the query vector q i , the query index name, the distribution distance distance between the query vector q i and the retrieval target set V, and the number h of out-of-distribution queries; the number h of out-of-distribution queries represents the number of queries where the query vector is inconsistent with the distribution of the retrieval target set V;

[0011] The distribution awareness module is used to: calculate the distribution distance distance between the query vector q i and the retrieval target set V, and update the number h of out-of-distribution queries; determine whether the number h of out-of-distribution queries is greater than the threshold threshold. If it is less than or equal to the threshold threshold, use the log recording module to record the query log. If it is greater than the threshold threshold, from the log recording module, obtain the corresponding query vector q i in the log recording module as t i , form a historical query set T = {t 1 , t 2 … t l}, and use the distribution detection module to perform the following processing:

[0012] The distribution detection module is used to calculate the distribution difference coefficient differ between the historical query set T and the retrieval target set V, and determine whether it is necessary to adjust the distribution of the retrieval target set V through the historical query set T. When the difference coefficient differ is greater than 0, the distribution mapping module constructs a distribution optimization index, updates the number h of out-of-distribution queries to the initial value 0, and updates the distribution interval of the retrieval target set V according to the historical query set T; when the difference coefficient differ is less than or equal to 0, the existing index is not adjusted.

[0013] The distribution mapping module is used to construct a distribution optimization index for the retrieval target set V based on the historical query set T and the retrieval target set V;

[0014] The construction of the distribution optimization index includes the following steps:

[0015] a. Brute-force search for each vector v in the retrieval target set V i for the j exact nearest neighbors {z 1 , z 2 … z j} in the historical query set T, and calculate the mean vector

[0016] b. Calculate the mapping set F = {f 1 , f 2 … f n} of the retrieval target set V, where f i in the mapping set F corresponds one-to-one to v i in the retrieval target set V;

[0017] The calculation formula for the mapping set F of the retrieval target set V is as follows:

[0018] F = V + βZ

[0019] where Z = {z 1avg , z 2avg … z navg} is composed of the mean vector z iavg ; β is an adjustable fusion coefficient;

[0020] c. Input the mapping set F into the index construction module to construct an index. After the construction is completed, replace the data part in the index with the retrieval target set V, that is, replace F = {f 1 , f 2 … f n} with V = {v 1 , v 2 … v n}, where f i corresponds to v iOne-to-one correspondence; at the same time, retain the structural information of the index, where the structural information includes the distribution information of the retrieval target set V and the historical query set T, and reduce the retrieval steps between the retrieval target vectors in the index; use the log recording module to record query logs;

[0021] Preferably, the vector database system based on query distribution awareness further includes a storage module, an interface proxy module, and a graphical management interface; the storage module is used for persistent storage of the retrieval target set V, index data, and query logs; the graphical management interface provides an operable user interface for visual operation of the vector database; the interface proxy module is used to receive requests from the graphical management interface and forward them to the corresponding modules for execution;

[0022] Preferably, in the interval measurement module, the distance d Mi is calculated as follows:

[0023]

[0024] where μ represents the mean vector of the retrieval target set V, and S is the covariance matrix of the retrieval target set V;

[0025] The distribution interval of the retrieval target set V is denoted as [low, high]; low represents the lower limit of the distribution offset of the retrieval target set V, and its calculation method is as follows:

[0026] low = quantile(U(d M ), percent)

[0027] high represents the upper limit of the distribution offset of the retrieval target set V, and its calculation method is as follows:

[0028] high = quantile(U(d M ), 1 - percent)

[0029] where quantile(set, percent) represents the value at the percent-th percentile after arranging the elements in the set in ascending order;

[0030] Preferably, the calculation method of the distribution distance distance is as follows:

[0031]

[0032] where μ represents the mean vector of the retrieval target set V, and S is the covariance matrix of the retrieval target set V;

[0033] The calculation method of the number h of out-of-distribution update queries includes: the initial value of h is 0, and when the distribution distance distance is less than the distribution interval low or greater than the distribution interval high, h is incremented by 1;

[0034] Preferably, the calculation method of the distribution difference coefficient differ includes the following steps:

[0035] S1. Sample two non-overlapping sets V 1 and V 2 from the retrieval target set V, where each set contains k vectors, and calculate the Wasserstein distance d 1 between set V 2 and set V w (V 1 , V 2 );

[0036] S2. Randomly sample a set T 1 from the historical query set T, where the set T 1 contains k vectors, calculate the Wasserstein distance d 1 between set T 1 and set V w (T 1 , V 1 ), and calculate the Wasserstein distance d 1 between set T 2 and set V w (T 1 , V 2 );

[0037] S3. The calculation formula of the difference coefficient differ is as follows:

[0038]

[0039] where α is an adjustable coefficient.

[0040] Compared with the prior art, the present invention has the following beneficial effects: the present invention proposes a vector database system based on query distribution perception, which is composed of an index construction module, a vector retrieval module, an interval measurement module, a log recording module, a distribution perception module, a distribution detection module, a distribution mapping module, a storage module, an interface proxy module and a graphical management interface. Through the index construction, vector retrieval and storage modules, the problems of vector similarity calculation and data persistence storage are solved; through the log recording module and the distribution perception module, the user's query history is recorded, the distribution changes of user queries are timely perceived, and the index structure in the vector database is dynamically adjusted to ensure the retrieval quality of user queries with the same distribution and different distributions; through the distribution detection module, the difference coefficient is used to intuitively measure the distribution of user historical query records and vectors in the library; through the distribution mapping module, the distribution of the retrieval target set and the user's historical query records is integrated, which effectively reduces the search path during retrieval, improves the retrieval efficiency and recall rate, and does not need to change the index construction method.

[0041] However, existing vector databases are usually based on the assumption that query vectors and vectors in the database follow the same distribution. Once the query vector and the database vector distribution are different, the query performance will drop sharply. The present invention provides an efficient solution for the storage and retrieval of vector data. At the same time, it monitors the distribution of user queries in real time. On the basis of ensuring the quality of queries of users with the same distribution, it effectively solves the problems of high latency and low recall rate caused by queries of users with different distributions, and ensures the stability and reliability of query quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Some specific embodiments of the present invention will be described in detail in an exemplary but not restrictive manner with reference to the drawings. The same reference numerals in the drawings indicate the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0043] Figure 1 It is a system architecture diagram of the present invention.

[0044] Figure 2 It is an execution flow chart of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0046] To make the technical solutions and advantages of the present invention more clear, the present invention will be further described below in conjunction with the accompanying drawings. Figure 1 is the system architecture diagram of the present invention, Figure 2 and Figure 2 is the execution flowchart of the present invention. Embodiments of the present invention are as follows:

[0047] An embodiment of the present invention provides a vector database system based on query distribution awareness, including an index construction module, a vector retrieval module, an interval measurement module, a log recording module, a distribution awareness module, a distribution detection module, a distribution mapping module, a storage module, an interface proxy module, and a graphical management interface;

[0048] The index construction module is used to construct an index for the retrieval target set V = {v 1 , v 2 … v n}, record the index name, where n is the number of vectors in the retrieval target set V; the retrieval target set V is composed of the base set in the Laion dataset; the index name is input by the user;

[0049] The method for constructing the index includes: a graph-based method, a tree-based method, and a quantization-based method, supported by the Faiss and Annoy open-source libraries;

[0050] The vector retrieval module is used to receive the query dataset Q = {q 1 , q 2 … q m} and the query index name, where m is the number of vectors in the query dataset Q; the query dataset Q is composed of the query set in the Laion dataset, belonging to different modalities from the retrieval target set V, with a certain distribution difference; for each element in the query dataset Q, query the index according to the query index name to obtain the top-k approximate nearest neighbor retrieval results;

[0051] The interval measurement module is used to traverse each vector v i in the retrieval target set V, calculate the distance d i between the vector v Mi and the retrieval target set V, form a distance set U(d M ), and after the traversal is completed, record the distribution interval [low, high] of the retrieval target set V, indicating the distribution offset degree of the retrieval target set V;

[0052] The calculation method of the distance d Mi is as follows:

[0053]

[0054] where μ represents the mean vector of the retrieval target set V, and S is the covariance matrix of the retrieval target set V; the distance d Mi reflects the deviation degree of the distribution of the vector v i in the retrieval target set V;

[0055] In the distribution interval [low, high], low represents the lower limit of the distribution deviation of the retrieval target set V, and its calculation method is as follows:

[0056] low = quantile(U(d M ), 0.05)

[0057] In the distribution interval [low, high], high represents the upper limit of the distribution deviation of the retrieval target set V, and its calculation method is as follows:

[0058] high = quantile(U(d M ), 1 - 0.05)

[0059] where quantile(set, percent) represents the value at the percent-th percentile after arranging the elements in the set in ascending order;

[0060] The log recording module is used to record query logs, and the query logs include the query vector q i 、the query index name, the query vector q i and the distribution distance distance between the query vector q and the retrieval target set V, and the number h of out-of-distribution queries; the number h of out-of-distribution queries represents the number of queries where the query vector is inconsistent with the distribution of the retrieval target set V;

[0061] The distribution awareness module is used to calculate the distribution distance distance between the query vector q i and the retrieval target set V, update the number h of out-of-distribution queries,

[0062] The calculation method of the distribution distance distance is as follows:

[0063]

[0064] where μ represents the mean vector of the retrieval target set V, and S is the covariance matrix of the retrieval target set V;

[0065] The calculation method for updating the number h of out-of-distribution queries is: the initial value of h is 0, and when the distribution distance distance is less than the distribution interval low or greater than the distribution interval high, h is incremented by 1;

[0066] Determine whether the number h of out-of-distribution queries is greater than the threshold threshold. In this embodiment, the value of threshold is n * 0.3. If it is less than or equal to the threshold threshold, use the log recording module to record the query log. If it is greater than the threshold threshold, from the log recording module, obtain the corresponding query vector q in the log recording module according to the query index name i as t i , to form the historical query set T = {t 1 , t 2 … t l}, and use the distribution detection module to perform the following processing:

[0067] The distribution detection module is used to calculate the distribution difference coefficient differ between the historical query set T and the retrieval target set V, and determine whether it is necessary to adjust the distribution of the retrieval target set V through the historical query set T. When the difference coefficient differ is greater than 0, the distribution mapping module constructs a distribution optimization index, updates the number h of out-of-distribution queries to the initial value 0, and updates the distribution interval [low, high] according to the historical query set T; when the difference coefficient differ is less than or equal to 0, do not adjust the existing index;

[0068] The calculation method of the distribution difference coefficient includes the following steps:

[0069] S1, sample two non-overlapping sets V 1 、V 2 from the retrieval target set V, each set contains k vectors, and calculate the Wasserstein distance d 1 between set V 2 and set V w (V 1 , V 2 );

[0070] S2, randomly sample a set T 1 from the historical query set T, the set T 1 contains k vectors, calculate the Wasserstein distance d 1 between set T 1 and set V w (T 1 , V 1 ), and calculate the Wasserstein distance d 1 between set T 2 and set V w (T 1 , V 2 );

[0071] S3, the calculation formula of the difference coefficient differ is as follows:

[0072]

[0073] Among them, α is an adjustable coefficient, and in this embodiment, α takes the value of 0.125; the Wasserstein distance can effectively capture the spatial distribution relationship between vector sets, so as to judge the distribution difference between vector sets;

[0074] The distribution mapping module is used to construct a distribution optimization index for the retrieval target set V based on the historical query set T and the retrieval target set V;

[0075] The construction of the distribution optimization index includes the following steps:

[0076] a. Brute-force search for each vector v in the retrieval target set V i The j exact nearest neighbors {z 1 , z 2 … z j} in the historical query set T, and calculate the mean vector In this embodiment, j takes the value of 7; the mean vector z iavg Conforms to the distribution characteristics of the historical query set T, and has a similar meaning to the vector v i To represent the distribution of the vector v i In the historical query set T;

[0077] b. Calculate the mapping set F of the retrieval target set V = {f 1 , f 2 … f n}, and the f i in the mapping set F corresponds one-to-one with the v i calculated in the retrieval target set V;

[0078] The calculation formula for the mapping set F of the retrieval target set V is as follows:

[0079] F = V + βZ

[0080] where Z = {z 1avg , z 2avg … z navg}, which is composed of the mean vector z iavg , the distribution of the set Z is consistent with the historical query set T, and contains the vector information in the retrieval target set V; β is an adjustable fusion coefficient for adjusting the weight, and in this embodiment, it takes the value of 0.67; the mapping set F contains the distribution characteristics of the retrieval target set V and the historical query set T. For the top-k retrieval results in the retrieval target set V returned for each query in the historical query set T, the distance between the corresponding mapping vectors of the retrieval results is closer, making it easier to connect with each other when constructing the index, thereby improving the retrieval efficiency and recall rate;

[0081] c. Input the mapping set F into the index construction module to construct an index. After the construction is completed, replace the data part in the index with the retrieval target set V from the mapping set F, that is, replace F = {f 1 , f 2 … f n} with V = {v 1 , v 2 … v n}, where f i corresponds to v i one by one; at the same time, retain the structural information of the index. The structural information includes the distribution information of the retrieval target set V and the historical query set T, reducing the retrieval steps between the retrieval target vectors in the index. Since the mapping vector cannot replace the original vector for retrieval distance and returning retrieval results, it is necessary to replace it with the original vector; use the log recording module to record query logs;

[0082] The storage module is used for persistent storage of the retrieval target set V, index data, and query logs, and stores them in the form of files;

[0083] The graphical management interface provides an operable user interface for visualizing and operating the vector database;

[0084] The interface proxy module is used to receive requests from the graphical management interface and forward them to the corresponding modules for execution;

[0085] As described above, only some specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A vector database system based on query distribution awareness, characterized in that: It includes index building module, vector retrieval module, interval measurement module, log recording module, distribution perception module, distribution detection module and distribution mapping module; The index building module is used to: retrieve the target set V = {v1, v2...v n }Build an index and record the index name, where n is the number of vectors in the retrieval target set V; The vector retrieval module is used to: receive a query data set Q={q1,q2…q m } and the query index name, m is the number of vectors in the query dataset Q; for each element in the query dataset Q, query the index according to the query index name to obtain the top-k approximate nearest neighbor retrieval results; The interval measurement module is used to traverse each vector v in the retrieval target set V i , calculate the vector v i The distance d from the retrieval target set V Mi , forming a distance set U(d M ), after the traversal is completed, the distribution interval of the retrieval target set V is recorded, indicating the distribution deviation degree of the retrieval target set V; The logging module is used to record query logs, including query vectors q i , query index name, query vector q i The distribution distance distance from the retrieval target set V and the number of out-of-distribution queries h; the number of out-of-distribution queries h represents the number of queries whose query vectors are inconsistent with the distribution of the retrieval target set V; The distribution perception module is used to calculate the query vector q i The distribution distance distance of the retrieval target set V is updated, and the number of out-of-distribution queries h is updated; it is determined whether the number of out-of-distribution queries h is greater than the threshold threshold. If it is less than or equal to the threshold threshold, the log recording module is used to record the query log. If it is greater than the threshold threshold, the corresponding query vector q in the log recording module is obtained from the log recording module according to the query index name. i As t i , forming a historical query set T = {t1, t2…t l }, use the distribution detection module to perform the following processing: The distribution detection module is used to calculate the distribution difference coefficient differ between the historical query set T and the retrieval target set V, and determine whether it is necessary to adjust the distribution of the retrieval target set V through the historical query set T. When the difference coefficient differ is greater than 0, the distribution mapping module constructs a distribution optimization index, updates the number of out-of-distribution queries h to an initial value of 0, and updates the distribution interval of the retrieval target set V according to the historical query set T; when the difference coefficient differ is less than or equal to 0, the existing index is not adjusted; The distribution mapping module is used to construct a distribution optimization index for the retrieval target set V based on the historical query set T and the retrieval target set V; The construction of the distribution optimization index comprises the following steps: a, brute force search to retrieve each vector v in the target set V i The j exact nearest neighbors {z1,z2…z j }, calculate the mean vector b. Calculate the mapping set F of the retrieval target set V = {f1,f2…f n }, f in the mapping set F i And calculate the retrieval target set V in v i One to one correspondence; The calculation formula for calculating the mapping set F of the retrieval target set V is as follows: F=V+βZ Where Z = {z 1avg ,z 2avg …z navg }, by the mean vector z iavg composition; β is the adjustable fusion coefficient; c. Input the mapping set F into the index building module to build the index. After the construction is completed, replace the data part in the index from the mapping set F to the retrieval target set V; at the same time, retain the structural information of the index, the structural information includes the distribution information of the retrieval target set V and the historical query set T, and shorten the retrieval steps between the retrieval target vectors in the index; use the logging module to record the query log.

2. A query distribution-aware vector database system according to claim 1, characterized in that: It also includes a storage module, an interface proxy module, and a graphical management interface; The storage module is used to persistently store the retrieval target set V, index data, and query logs; The graphical management interface provides an operable user interface to visualize the operation vector database; The interface proxy module is used to receive a graphical management interface request and forward it to a corresponding module for execution.

3. A query distribution-aware vector database system according to claim 1, characterized in that: In the interval measurement module, the distance d Mi The calculation method is as follows: Among them, μ represents the mean vector of the retrieval target set V, and S is the covariance matrix of the retrieval target set V; The distribution interval of the retrieval target set V is recorded as [low, high]; low represents the lower limit of the distribution offset of the retrieval target set V, and its calculation method is as follows: low=quantile(U(d M ),percent) High represents the upper limit of the distribution deviation of the retrieval target set V, which is calculated as follows: high=quantile(U(d M ),1-percent) Among them, quantile(set,percent) represents the percentile value of the elements in the set after they are arranged from small to large.

4. A query distribution-aware vector database system according to claim 3, characterized in that: The calculation method of the distribution distance is as follows: Among them, μ represents the mean vector of the retrieval target set V, and S is the covariance matrix of the retrieval target set V; The calculation method for updating the number of out-of-distribution queries h includes: the initial value of h is 0, and when the distribution distance distance is less than the distribution interval low or greater than the distribution interval high, h is increased by 1.

5. A query distribution-aware vector database system according to claim 4, characterized in that: The calculation method of the distribution difference coefficient differ includes the following steps: S1, sample two disjoint sets V1 and V2 from the retrieval target set V, each of which contains k vectors, and calculate the Wasserstein distance d between set V1 and set V2 w (V1,V2); S2, randomly sample a set T1 from the historical query set T, the set T1 contains k vectors, and calculate the Wasserstein distance d between the set T1 and the set V1 w (T1, V1), calculate the Wasserstein distance d between set T1 and set V2 w (T1, V2); S3, the calculation formula of the difference coefficient differ is as follows: Among them, α is an adjustable coefficient.

Citation Information

Patent Citations

  • Method and device for face image retrieval, computer device and storage medium

    CN108932321A

  • Hybrid vector retrieval method and device for high-concurrency scene

    CN116166690A

  • System log anomaly detection system and method based on comparative learning and attention mechanism

    CN118585367A

  • Framework for continuously optimizing graph structure in high-dimensional approximate nearest neighbor search

    CN118885629A

  • System and method for the indexing and retrieval of semantically annotated data using an ontology-based information retrieval model

    US20160179945A1