Selective data collection from remote devices

By determining the relationship between confidence level and cost rating in a distributed system and selecting appropriate remote device samples, the problem of inefficient data acquisition of remote devices is solved, and efficient and economical data acquisition effect is achieved.

CN120067168APending Publication Date: 2025-05-30FORD GLOBAL TECH LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411706356.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In distributed systems, remote devices such as vehicles are difficult to effectively collect and transmit large amounts of data due to bandwidth and storage limitations, resulting in inefficient data acquisition.

Method used

The relationship between confidence level and cost rating is determined by computer, the input of confidence value and cost value is received, the appropriate remote device sample is selected, and the data request is transmitted to the remote device in the sample to achieve efficient utilization of resources.

Benefits of technology

It realizes the saving of memory and network bandwidth resources while ensuring the accuracy of data acquisition, and improves the efficiency and cost-effectiveness of data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067168A_ABST
    Figure CN120067168A_ABST
Patent Text Reader

Abstract

The present disclosure provides for selective data collection from remote devices. A computer is programmed to determine a relationship between a confidence level and a cost rating for collecting a set of data from a plurality of remote devices, then receive an input selecting a confidence value for the confidence level and a cost value for the cost rating, select an actual sample of the remote devices according to the confidence value, and transmit the actual sample to the remote devices. And transmitting a request for the set of data to the remote device in the actual sample. The confidence level indicates a statistical confidence of the set of data based on candidate samples of the remote device. The cost rating indicates a cost of collecting the set of data from the candidate samples of the remote device. The confidence value and the cost value coincide with a relationship between the confidence level and the cost rating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a system for selectively collecting data from remote devices. Background Art

[0002] A distributed system is a computing system in which its components are located on different networked computers. The networked computers communicate and coordinate their actions by passing messages to each other. The networked computers are independent of each other, and each networked computer has its own local memory. Summary of the Invention

[0003] The system described herein provides a resource - efficient way to capture data from remote devices (such as vehicles), which are part of a distributed system with a central computer. Distributed systems may face challenges in efficiently transmitting relevant data. Vehicles and some other types of remote devices can generate large amounts of data, and due to limitations such as available memory and / or computer network bandwidth, it may be impractical and / or technically infeasible to collect (including transmit and store) all of this data. For example, different remote devices can have different per - device bandwidth limitations. In the system described herein, a computer is programmed to determine the relationship between a confidence level and a cost rating for collecting a set of data from multiple remote devices, then receive an input of a confidence value for selecting the confidence level and a cost value for the cost rating, select an actual sample of the remote devices according to the confidence value, and transmit a request for the set of data to the remote devices in the actual sample. The confidence level indicates the statistical confidence of the set of data based on candidate samples of the remote devices. The cost rating indicates the cost of collecting the set of data from candidate samples of the remote devices, i.e., the technical effort in terms of processor cycles, bandwidth usage, etc. The confidence value and the cost value are consistent with the relationship between the confidence level and the cost rating. For example, the input can include a number of the confidence value from which the cost value is determined according to the relationship, or the input can include a number of the cost value from which the confidence value is determined according to the relationship. Thus, the system saves resources such as available memory and network bandwidth while still capturing the data that the user is most interested in.

[0004] A computer includes a processor and a memory, and the memory stores instructions executable by the processor to perform the following operations: determine the relationship between a confidence level and a cost rating for collecting a set of data from multiple remote devices, then receive an input selecting a confidence value for the confidence level and a cost value for the cost rating, select an actual sample of the remote devices based on the confidence value, and transmit a request for the set of data to the remote devices in the actual sample. The confidence level indicates the statistical confidence of the set of data based on candidate samples of the remote devices. The cost rating indicates the cost of collecting the set of data from the candidate samples of the remote devices. The confidence value and the cost value are consistent with the relationship between the confidence level and the cost rating.

[0005] In one example, the remote device can be a vehicle.

[0006] In one example, the set of data can be defined by a set of characteristics of the data. In additional examples, the characteristics can include the model of the remote device.

[0007] In yet another example, the characteristics can include the geographical region containing the remote device.

[0008] In yet another example, the characteristics can include the usage of features of the remote device. In yet another example, the instructions can further include instructions for performing the following operations: select multiple classifications of data transmitted within the remote device to be included in the set of data according to a mapping associating the classifications with the features. In yet another example, the instructions can further include instructions for performing the following operations: determine a classification name for at least one of the classifications, the classification name being specific to the model of the remote device; and include the classification name in the request for the remote device.

[0009] In yet another example, the feature can be a first feature; the classification can be a first classification; the mapping can associate a second classification with a second feature; and the instructions can further include instructions for performing the following operations: output a prompt to the user to suggest that the user add the second feature to the set of characteristics in response to an overlap between the first classification and the second classification.

[0010] In yet another example, the feature can be a first feature, and the instructions can further include instructions for performing the following operations: output a prompt to the user to suggest that the user add a second feature to the set of characteristics in response to a correlation between a request for the first feature and a request for the second feature.

[0011] In yet another example, the characteristics can include events affecting the remote device.

[0012] In another example, the input can be a second input, and the instructions can further include instructions for receiving a first input that specifies the characteristic.

[0013] In another example, the instructions can further include instructions for: determining a population of the remote devices based on the characteristic, and randomly selecting an actual sample from the population. In yet another example, the characteristic can be a first characteristic; the actual sample can be a first actual sample; and the instructions can further include instructions for: selecting a second actual sample of the remote devices based on a second characteristic related to the first characteristic.

[0014] In one example, the instructions can further include instructions for: receiving an input that specifies a maximum value of the cost rating, and determining a candidate confidence value based on a relationship between the confidence level and the cost rating.

[0015] In one example, the instructions can further include instructions for: receiving an input that specifies a minimum value of the confidence level, and determining a candidate cost value based on a relationship between the confidence level and the cost rating.

[0016] In one example, the instructions can further include instructions for: selecting the actual sample of the remote devices based on a bandwidth limitation of a corresponding remote device in the actual sample.

[0017] In one example, the cost rating can be a function of the number of remote devices in the actual sample.

[0018] In one example, the confidence level can be a function of the number of remote devices in the actual sample.

[0019] A method includes determining a relationship between a confidence level and a cost rating for collecting a set of data from a plurality of remote devices, then receiving inputs of a confidence value for selecting the confidence level and a cost value for the cost rating, selecting an actual sample of the remote devices based on the confidence value, and transmitting a request for the set of data to the remote devices in the actual sample. The confidence level indicates a statistical confidence of the set of data based on candidate samples of the remote devices. The cost rating indicates a cost of collecting the set of data from the candidate samples of the remote devices. The confidence value and the cost value are consistent with the relationship between the confidence level and the cost rating. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a block diagram of an exemplary communication system between a computer and remote devices.

[0021] Figure 2 It is a diagram of the mapping between the characteristics of a remote device and the classification names of data transmissions related to the characteristics.

[0022] Figure 3 It is a diagram of an exemplary sampling of a remote device.

[0023] Figure 4 It is a flowchart of an exemplary process for collecting data from a remote device. Detailed Description

[0024] Referring to the accompanying drawings, in which like numerals indicate like parts throughout several views, computer 100 includes a processor and a memory, and the memory stores instructions executable by the processor to perform the following operations: determining the relationship between a confidence level and a cost rating for collecting a set of data from a plurality of remote devices 105, then receiving an input selecting a confidence value for the confidence level and a cost value for the cost rating, selecting actual samples 305, 310 of the remote devices 105 based on the confidence value, and transmitting a request for the set of data to the remote devices 105 among the actual samples 305, 310. The confidence level indicates the statistical confidence of the set of data based on candidate samples of the remote device 105. The cost rating indicates the cost of collecting the set of data from the candidate samples of the remote device 105. The confidence value and the cost value are consistent with the relationship between the confidence level and the cost rating.

[0025] Referring to Figure 1 , computer 100 is a microprocessor-based computing device, such as a general-purpose computing device including a processor and a memory. The memory of computer 100 may include media for storing instructions executable by the processor and for electronically storing data and / or databases, and / or computer 100 may include a structure such as the foregoing structure providing programming. Computer 100 may be a plurality of computers coupled together.

[0026] Computer 100 may communicate with remote devices 105 via network 110. Network 110 represents one or more mechanisms by which computer 100 may communicate. Thus, network 110 may be one or more of various wired or wireless communication mechanisms, including any desired combination of wired (e.g., cable and fiber) and / or wireless (e.g., cellular, wireless, satellite, microwave, and radio frequency) communication mechanisms, and any desired network topology (or multiple topologies when utilizing multiple communication mechanisms). Exemplary communication networks include wireless communication networks providing data communication services (e.g., using Bluetooth, IEEE 802.11, etc.), local area networks (LANs), and / or wide area networks (WANs), including the Internet.

[0027] The remote device 105 is or includes a computing device. The remote device 105 can be mobile, i.e., capable of physically moving around or being moved around during operation. For example, the remote device 105 can be a vehicle, such as Figure 1 depicted. The vehicle can be any passenger or commercial vehicle, such as a sedan, truck, sport utility vehicle, crossover vehicle, van, minivan, taxi, bus, etc. As another example, the remote device 105 can be a portable computing device, such as a cellular phone, a tablet computer, a wearable computing device such as a smartwatch, etc. The techniques described herein are particularly useful for collecting data from a mobile remote device 105 with potentially limited bandwidth, such as a vehicle.

[0028] The remote device 105 may have a bandwidth limitation. A bandwidth limitation is an upper bound on the amount or rate of data transmitted by the remote device 105 to the network 110. The bandwidth limitation can be specific to the remote device 105; i.e., different remote devices 105 can have different bandwidth limitations. The bandwidth limitation can be specified by the model of the remote device 105; for example, some models of the remote device 105 can have a bandwidth limitation of 12 MB per month, and other models can have a bandwidth limitation of 30 MB per month.

[0029] Referring Figure 2 , the remote device 105 can generate data due to normal operation (i.e., not in response to a request from the computer 100). The data includes data transmitted within the remote device 105, e.g., data sent over the communication network or bus of the remote device 105. The data can include structured data and unstructured data. Structured data is data organized in a standardized format. For example, data sent over a vehicle's controller area network (CAN) bus typically adopts a database container (.dbc) file format, which is a structured data type. Some sensors, such as cameras, can produce unstructured data. The data can be time series data. As is commonly understood and for the purposes of this disclosure, time series data is the value of one or more variables at discrete consecutive time points.

[0030] As described below, computer 100 determines a set of data to be collected from remote device 105, and the set of data is a subset of all the data generated on remote device 105. The set of data is defined by a set of characteristics of the data. For the purposes of this disclosure, a "characteristic" of data is a fact that is true for some of the data generated by remote device 105. Thus, each characteristic defines a subset of the data generated by remote device 105, i.e., the subset for which the characteristic holds or is true. For example, characteristics can include features 205 of remote device 105, classification 210 of the data, classification name 215 of the data, model of remote device 105, geographical region in which remote device 105 may be located, environmental conditions around remote device 105 when the data is generated, time and / or date when the data is generated, and / or events affecting remote device 105 when the data is generated, each of which will be described in turn below. The characteristics can also include other types of facts, such as demographic data (such as age) about the owner or operator of remote device 105, data describing the quality of remote device 105, etc.

[0031] The characteristics can include the use of features 205 of remote device 105. Feature 205 is a certain ability of remote device 105. As an example of a vehicle, feature 205 can include an Advanced Driver Assistance System (ADAS). ADAS is an electronic technology that assists a driver in achieving driving functions and parking functions. Examples of ADAS include forward proximity detection, lane departure detection, blind spot detection, brake actuation, adaptive cruise control, and lane keeping assistance systems. As another example, feature 205 can include basic operations of the vehicle, such as braking, accelerating, steering, etc. As another example, feature 205 can include vehicle monitoring, such as brake prediction (e.g., life cycle estimation of brakes), Diagnostic Trouble Codes (DTCs), etc. Other examples of feature 205 include climate control for the passenger compartment of the vehicle, infotainment options, and the status and operation of vehicle components (such as lights, latches and locks, sensors, etc.).

[0032] Characteristics can include classification 210 of the data. Classification 210 is the type of data, e.g., the type of data transmitted within remote device 105, e.g., the type of payload in a message sent over a vehicle's communication network or bus. For example, classification 210 can include types of sensor data, such as vehicle speed, sonar, ambient temperature, temperature of the passenger compartment, etc. Classification 210 can also include commands to components, such as braking force, fan speed, etc.

[0033] The computer 100 may store a first mapping 220 that associates the classification 210 with the feature 205. The first mapping 220 includes a plurality of first associations 225 between the feature 205 and the classification 210. Each first association 225 between the feature 205 and the classification 210 indicates that the feature 205 uses the classification 210, i.e., the execution or operation of the feature 205 is based on the classification 210. For example, brake prediction (feature 205) is determined based on data including vehicle speed and braking force over time (classification 210) to estimate the life cycle of the brakes. The computer 100 may be programmed to determine the classification 210 of a given feature 205 according to the first mapping 220 (i.e., by looking up the first mapping 220). The computer 100 may be programmed to select a plurality of classifications 210 of the data transmitted within the remote device 105 according to the first mapping 220 to include in the set of data, e.g., by selecting the classifications 210 determined to be associated with the feature 205 of interest (e.g., the feature 205 identified by an input from a user). The first mapping 220 may be stored in the memory of the computer 100, for example, as a set of feature-classification pairs (e.g., {(feature A, classification A), (feature A, classification B),...}) or as a matrix having columns for each feature 205 and rows for each classification 210 (or vice versa), where a binary variable in each entry indicates whether the feature 205 of the column is associated or not associated with the classification 210 of the row.

[0034] The classification 210 has a classification name 215. The classification name 215 is an identifier used within the remote device 105 to refer to the classification 210. The classification name 215 may be specific to one or more models of the remote device 105. In other words, the classification name 215 of a particular classification 210 may be different on different models of the remote device 105. For example, vehicle speed (classification 210) (i.e., how fast the vehicle is traveling) may be identified as "Veh_Eng_Actl" (classification name 215) on some vehicle models and as "Veh_Spd" (classification name 215) on other vehicle models.

[0035] The computer 100 may store a second mapping 230 that associates classification names 215 with classifications 210. The second mapping 230 includes a plurality of second associations 235 between the classifications 210 and the classification names 215. Each second association 235 between a classification 210 and a classification name 215 indicates that the classification name 215 refers to the classification 210 on a certain model of the remote device 105. The computer 100 may be programmed to determine the classification name 215 of a given classification 210 and the model of the remote device 105 according to the second mapping 230 (i.e., by looking up the second mapping 230). The second mapping 230 may be stored, for example, as a matrix in the memory of the computer 100 having columns for each classification 210 and rows for each model of the remote device 105 (or vice versa), where each entry includes the classification name 215 of the classification 210 in the column used in the model of the remote device 105 in that row.

[0036] The characteristics may include the model of the remote device 105. The model of the remote device 105 is a specific product design or version embodied in the context of the manufacturer's product range or series.

[0037] For example, when the remote device 105 generates data of interest, the characteristics may include the geographical area in which the remote device 105 is located. As an example of a vehicle, the geographical area may be specified by the type of road on which the vehicle travels (e.g., restricted access highway, standard highway, urban street, residential street, gravel road, parking lot, etc.), or by the identification of the road on which the vehicle travels (e.g., Route 66). As another example, the geographical area may be specified by municipal or jurisdictional boundaries; by population density (e.g., urban, suburban, rural); by state or country; by some combination of the foregoing; and so on.

[0038] The characteristics may include one or more environmental conditions that the remote device 105 experiences when generating data of interest. Environmental conditions may include traffic density, road surface type (e.g., paved or gravel), weather, etc. Weather may include precipitation (e.g., rain, snow, no precipitation, etc.), cloud cover (e.g., clear, partly cloudy, cloudy, etc.), visibility (e.g., foggy, hazy, clear, etc.), ambient temperature, etc. The remote device 105 may use built-in sensors (such as vibration sensors for road surface type, rain sensors, temperature sensors, light sensors, etc.) to determine the environmental conditions. Alternatively or additionally, the remote device 105 may receive external data indicating environmental conditions (such as traffic density and weather) from the network 110.

[0039] The characteristics may include the time at which the remote device 105 generates data of interest. The time of interest may be specified using the time of day, the day of the week, and / or a range of dates. For example, the data of interest may be data generated during peak hours, which are specified as a particular period during a weekday.

[0040] The characteristics may include events that affect the remote device 105 when the data is generated. For example, an event may be defined when a certain quantity is above or below a threshold or within or outside a range. For example, an event may occur when the braking force or deceleration is above a threshold.

[0041] Reference Figure 3 Referring, the computer 100 may be programmed to determine a population 300 of remote devices 105 based on a set of characteristics. The population 300 is a collection of remote devices 105 that have the set of characteristics. For example, if the characteristics are that the model is Model A and the feature 205 is automatic air refresh, the population 300 includes all remote devices 105 that are Model A and equipped with automatic air refresh, and the population 300 excludes other remote devices 105, e.g., remote devices 105 that are Model B or remote devices 105 that are Model A but lack automatic air refresh. The computer 100 may receive the set of characteristics as an input. The computer 100 may consult a certain list of characteristics of the particular remote device 105, e.g., a database of remote devices 105 associated with their respective characteristics. The remote device 105 may be identified in the database, etc., by a unique identifier (e.g., vehicle identification number (VIN)). The database may specify characteristics that represent or indicate static properties of the remote device 105, such as model, feature 205, classification 210, rather than characteristics that represent or indicate transient properties, such as geographical area, environmental conditions, time, events.

[0042] The computer 100 can be programmed to randomly select actual samples 305, 310 from the population 300 based on input from the user. The actual samples 305, 310 are the set of remote devices 105 from which data will be collected. As an overall overview, the computer 100 can determine the population 300 based on characteristics input by the user, as described above. The computer 100 determines the relationship between the confidence level and the cost rating for collecting the set of data from the remote devices 105. The set of data can be specified by input from the user (e.g., an input listing characteristics). The relationship describes a trade-off between the confidence level and the cost rating, i.e., a higher confidence value generally requires a higher cost value. The computer 100 can select a point on the relationship between the confidence level and the cost rating based on input from the user, the point being represented by a confidence value and a cost value. The input can include a confidence value or a cost value, and the computer 100 can determine the other of the confidence value and the cost value based on the relationship. Then, the computer 100 can randomly select the actual samples 305, 310 from the population 300 of remote devices 105 to achieve the selected confidence level, and transmit a request for the set of data to the remote devices 105 in the actual samples 305, 310.

[0043] The cost rating indicates the cost of collecting the set of data from the candidate samples of the remote devices 105. The cost represents the technical effort to obtain the set of data from the remote devices 105, such as processor cycles, bandwidth usage, etc. Different types of costs can be converted to a common unit to make different types of technical effort commensurate. The cost rating can be a function of the amount of the set of data to be collected, e.g., C = f(X), where C is the cost rating and X is the amount of data to be collected. The function can define a positive relationship between the data amount and the cost rating, i.e., increasing the data amount results in an increase in the cost rating. For example, the function can be a linear relationship between the cost rating and the amount of data to be collected, i.e., C = mX, where m is the cost per unit of data. As another example, the function can be a combination of different linear relationships for sub-populations of the population 300, e.g., as shown in the following formula in the case of two sub-populations:

[0044] C = (m 1 k + m 2 (1 - k))X

[0045] where m 1 and m 2are the respective per-unit data costs for the first and second sub-populations, and k is the proportion of population 300 in the first sub-population. The cost rating can be a function of the number of remote devices 105 in the sample, e.g., by virtue of the amount of data in the set to be collected and the amount of data to be collected from each remote device 105, e.g., X = Nx, where N is the number of remote devices 105 in the sample and x is the data collected from each remote device 105. The data x for each remote device 105 can be limited at the bandwidth limit.

[0046] The confidence level indicates the statistical confidence in the set of data based on the candidate sample of remote devices 105. The confidence level can be a function of the amount of data in the set to be collected, e.g., according to the Cochran formula known in statistics. The Cochran formula is given by the following expression:

[0047]

[0048] where Z is the z-value indicating the confidence level, p is the estimated proportion of data exhibiting the property of interest available for collection, and e is the desired margin of error. The confidence level can be a function of the number of remote devices 105 in the sample, e.g., by virtue of the amount of data in the set to be collected and the amount of data to be collected from each remote device 105, as described above for the cost rating.

[0049] Computer 100 is programmed to determine the relationship between the confidence level and the cost rating for collecting the set of data from population 300 of remote devices 105. For example, computer 100 can parametrically define the relationship for the amount of data to be collected by using the formulas described above for the cost rating as a function of the amount of data to be collected and for the confidence level as a function of the data to be collected. As another example, computer 100 can determine an expression that directly relates the cost rating and the confidence level by solving a system of equations using the above formulas. The relationship represents a trade-off between cost and confidence, where greater confidence generally requires higher cost.

[0050] The computer 100 can be programmed to determine a candidate cost value and a candidate confidence value based on the relationship between the confidence level and the cost rating according to the user's input. The input from the user can effectively select a point on the relationship, for example, by specifying a minimum value of the confidence level or a maximum value of the cost rating, where the other quantity is determined to be consistent with the relationship, such as by solving the above formula. For example, the computer 100 can receive an input specifying the minimum value of the confidence level, which is used as the candidate confidence value, and then the computer 100 can determine the candidate cost value based on the candidate confidence value consistent with the relationship. The candidate cost value can be the minimum cost value that at least provides the minimum value of the input of the confidence level. Again, for example, or during different executions or iterations, the computer 100 can receive an input specifying the maximum value of the cost rating, which is used as the candidate cost value, and then the computer 100 can determine the candidate confidence value based on the candidate cost value consistent with the relationship. The candidate confidence value can be the maximum confidence value achievable within the maximum value of the input of the cost rating.

[0051] The computer 100 can be programmed to provide the user with options to iterate the above steps with different inputs. The computer 100 can output the candidate cost value and the candidate confidence value, and then the user can input a command to cause the computer 100 to allow a new input. The user can provide a different set of characteristics, different minimum values of the confidence level, and / or different maximum values of the cost rating, and the computer 100 can thereby derive different candidate cost values and / or candidate confidence values. Once the user has completed the iteration, the values of the final candidate cost value and the final candidate confidence value are used as the cost value and the confidence value, respectively.

[0052] The computer 100 can be programmed to output a prompt to the user to suggest that the user add the feature 205 to the set of features. Then, the user can include the suggested feature 205 in the next iteration. For example, the computer 100 can output a prompt in response to an overlap between the classification 210 associated with the suggested feature 205 and the classification 210 associated with one of the features 205 already in the set of features. The computer 100 can determine that there is an overlap between two features 205 in response to the number or proportion of classifications 210 associated with the two features 205 exceeding a threshold. As another example, the computer 100 can output a prompt in response to a correlation between a request for the suggested feature 205 and a request for one of the features 205 already in the set of features. The correlation is a statistical relationship between the two features 205, e.g., the likelihood that the two features 205 have been selected together by past users. The computer 100 can output a prompt in response to the correlation exceeding a threshold. As yet another example, the computer 100 can output a prompt based on the output of a recommendation system executed by the computer 100. The recommendation system can be any suitable algorithm for recommending selections to the user based on selections made by past users, e.g., collaborative filtering, content-based filtering, hybrid filtering, etc.

[0053] The computer 100 is programmed to select a first actual sample 305 of the remote devices 105 based on a confidence value. For example, the computer 100 can select a plurality of remote devices 105 from the population 300 to provide a data volume sufficient to satisfy the Cochran's formula given the input confidence value. Selecting the first actual sample 305 can be based on the bandwidth limitations of the remote devices 105 in the first actual sample 305; e.g., lower bandwidth limitations require selecting a greater number of remote devices 105 to generate the same amount of data. The computer 100 can select the first actual sample 305 by randomly sampling the unique identifiers of the remote devices 105 in the population 300. Alternatively or additionally, the computer 100 can partition the population 300 into clusters according to characteristics and randomly sample from each cluster to ensure that all characteristics are well represented in the first actual sample 305.

[0054] The computer 100 can be programmed to select a second actual sample 305 of the remote devices 105 based on characteristics not included in the characteristics used to select the first actual sample 305. The additional characteristics can be related to the characteristics already included, e.g., a higher likelihood that the characteristics appear together in the same remote device 105. Thus, the second actual sample 305 can collect data on characteristics that are particularly cost-effective. The computer 100 can select the second actual sample 305 from the population 300 in the same manner as the first actual sample 305.

[0055] Figure 4is a flowchart showing an exemplary process 400 for collecting data from a remote device 105. The memory of computer 100 stores executable instructions for performing the steps of process 400, and / or programming may be implemented in a structure such as that described above. As an overall overview of process 400, computer 100 receives an input, determines a classification name 215, determines recommended features 205, receives additional selections of features 205, selects a compression method, outputs the relationship between a cost rating and the confidence level, and receives an input selecting the cost value and the confidence value. Computer 100 iterates the foregoing steps as long as the user desires to change the input. Next, computer 100 determines the definition of the set of data to include in the request, selects actual samples 305, 310 of remote device 105, transmits the request to remote device 105, and receives the requested data from remote device 105.

[0056] Process 400 begins at block 405, where computer 100 receives an input specifying a characteristic and an input specifying a maximum of a cost rating or a minimum of a confidence level.

[0057] Next, at block 410, computer 100 determines the classification name 215 covered by the characteristic received at block 405. Computer 100 determines classification 210 according to a first mapping 220 of features 205 included in the characteristic, and computer 100 determines the classification name 215 of those classifications 210 (and any other classifications 210 included in the input characteristic) according to a second mapping 230, as described above.

[0058] Next, at block 415, computer 100 outputs a prompt to the user to suggest that the user add recommended features 205, as described above.

[0059] Next, at block 420, computer 100 receives an input specifying which (if any) of the recommended features 205 to add to the characteristic.

[0060] Next, in block 425, computer 100 selects a compression method for remote device 105 to use when transmitting the requested data to computer 100. The compression method can be selected from a set of possible compression methods suitable for the data from remote device 105, e.g., no compression, compressive sensing, wavelets, principal component analysis (PCA), autoencoders, etc. For example, a user can provide an input to select the compression method. As another example, computer 100 can determine the compression method based on the characteristics received in block 405 (e.g., based on classification 210). Computer 100 can consult a lookup table that pairs classification 210 with compression methods. The pairing can be selected based on the data density of classification 210 and based on empirical testing of compression methods for different classifications 210. As an example, numerical time series data can be uncompressed, image data can be compressed with PCA, and so on.

[0061] Next, in block 430, computer 100 determines the relationship between the confidence level and the cost rating, as described above. The cost rating can depend on the compression method selected in block 425. For example, the cost rating can reduce the compression percentage of the selected compression method, e.g., C = (1–k)C unc , where k is the percentage reduction in data compression according to the compression method, and C unc is the cost rating of the uncompressed data. The cost rating C of the uncompressed data can be determined as described above for the cost rating. unc .

[0062] Next, in block 435, computer 100 receives an input of a confidence value for selecting a confidence level and a cost value for the cost rating. For example, the input can specify a minimum value for the confidence level, and computer 100 determines the cost value based on the minimum value, or the input can specify a maximum value for the cost rating, and computer 100 determines the confidence value based on the maximum value. In either case, the determination ensures that the confidence value and the cost value are consistent with the relationship between the confidence level and the cost rating.

[0063] Next, in decision block 440, computer 100 determines whether the user wants to iterate the determination of the confidence value and the cost value. For example, computer 100 can determine whether computer 100 has received an input that provides new characteristics or values or requests the provision of new characteristics or values. In response to an iteration request, process 400 returns to block 405 to receive new inputs. In response to the user being done, process 400 proceeds to block 445.

[0064] In block 445, computer 100 determines the definition of the data to request from remote device 105. The data can be defined according to the classification name 215 of the classification 210 selected from the final iteration and any classification 210 recommended for the second actual sample 305, as described above.

[0065] Next, in block 450, computer 100 determines population 300 based on characteristics, and computer 100 selects a first actual sample 305 and a second actual sample 305 of remote device 105 from population 300 by random sampling according to a confidence value, as described above.

[0066] Next, in block 455, computer 100 transmits a request for the set of data defined in block 445 to remote device 105 in the first actual sample 305 and the second actual sample 310 from block 450 via network 110. As described above, computer 100 includes classification name 215 in the request for remote device 105 according to second mapping 230 such that all remote device 105s can correctly interpret the request.

[0067] Next, in block 460, computer 100 receives the requested data from remote device 105 via network 110. After block 460, process 400 ends.

[0068] In general, the described computing systems and / or devices may employ any of a number of computer operating systems, including but not limited to the following versions and / or variants: Ford applications, AppLink / SmartDevice Link middleware, operating systems, Microsoft operating systems, Unix operating systems (e.g., the operating system released by Oracle Corporation of Redwood Shores, California), the AIX UNIX operating system released by International Business Machines Corporation of Armonk, New York, Linux operating systems, the Mac OSX and iOS operating systems released by Apple Inc. of Cupertino, California, the BlackBerry OS released by BlackBerry Limited of Waterloo, Canada, and the Android operating system developed by Google Inc. and the Open Handset Alliance, or the CAR Platform for infotainment provided by QNX Software Systems. Examples of computing devices include but are not limited to in-vehicle computers, computer workstations, servers, desktops, notebooks, laptop computers, or handheld computers, or some other computing system and / or device.

[0069] Computing devices typically include computer-executable instructions, where the instructions can be executed by one or more computing devices such as those listed above. The computer-executable instructions can be compiled or interpreted from computer programs created using a variety of programming languages and / or technologies, which alone or in combination include but are not limited to Java TM, C, C++, Matlab, Simulink, Stateflow, Visual Basic, Java Script, Python, Perl, HTML, etc. Some of these applications can be compiled and executed on virtual machines such as the Java virtual machine, the Dalvik virtual machine, etc. Generally, a processor (e.g., a microprocessor) receives instructions from, for example, a memory, a computer-readable medium, etc., and executes these instructions, thereby performing one or more processes, including one or more of the processes described herein. Such instructions and other data can be stored and transmitted using a variety of computer-readable media. Files in a computing device are typically a collection of data stored on a computer-readable medium such as a storage medium, random access memory, etc.

[0070] A computer-readable medium (also referred to as a processor-readable medium) includes any non-transitory (e.g., tangible) medium that participates in providing data (e.g., instructions) that can be read by a computer (e.g., by a processor of the computer). Such media can take many forms, including but not limited to non-volatile media and volatile media. Instructions can be transmitted via one or more transmission media, including optical fibers, wires, wireless communication, including internal components that make up a system bus coupled to a processor of the computer. Common forms of computer-readable media include, for example, RAM, PROM, EPROM, FLASH-EEPROM, any other memory chip or cartridge, or any other medium from which a computer can read.

[0071] The databases, data repositories, or other data stores described herein can include various mechanisms for storing, accessing / retrieving, and retrieving various data, including hierarchical databases, sets of files in a file system, application databases in a proprietary format, relational database management systems (RDBMSs), non-relational databases (NoSQL), graph databases (GDBs), etc. Each such data store is typically included within a computing device employing a computer operating system such as one of those mentioned above, and is accessed via a network in any one or more of a variety of ways. The file system can be accessed from the computer operating system and can include files stored in various formats. In addition to languages for creating, storing, editing, and executing stored programs such as the PL / SQL language mentioned above, RDBMSs typically also employ the Structured Query Language (SQL).

[0072] In some examples, system components may be implemented as computer-readable instructions (e.g., software) on one or more computing devices (e.g., servers, personal computers, etc.) and stored on a computer-readable medium associated therewith (e.g., disks, memories, etc.). A computer program product may include such instructions stored on a computer-readable medium for performing the functions described herein.

[0073] In the drawings, like reference numerals indicate like elements. Additionally, some or all of these elements may be altered. With respect to the media, processes, systems, methods, heuristics, etc. described herein, it should be understood that while the steps of such processes etc. have been described as occurring in a certain ordered sequence, such processes may be practiced by performing the steps in an order different from that described herein. It should also be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. The operations, systems, and methods described herein should always be implemented and / or performed in accordance with applicable owner / user manuals and / or safety guidelines.

[0074] The present disclosure has been described in an illustrative manner, and it should be understood that the terms used are of a descriptive nature and not restrictive. The adjectives "first" and "second" are used throughout this document as identifiers and are not intended to denote importance, order, or quantity. The use of "responsive to," "after determining...," etc. indicates a causal relationship and not merely a temporal relationship. Given the above teachings, many modifications and variations of the present disclosure are possible, and the present disclosure may be practiced in other ways than specifically described.

[0075] According to the present invention, there is provided a computer having a processor and a memory, the memory storing instructions executable by the processor to perform the following operations: determining a relationship between a confidence level and a cost rating for collecting a set of data from a plurality of remote devices, the confidence level indicating the statistical confidence of the set of data based on candidate samples of the remote devices, the cost rating indicating the cost of collecting the set of data from the candidate samples of the remote devices; then receiving an input selecting a confidence value of the confidence level and a cost value of the cost rating, the confidence value and the cost value being consistent with the relationship between the confidence level and the cost rating; selecting actual samples of the remote devices according to the confidence value; and transmitting a request for the set of data to the remote devices among the actual samples.

[0076] According to an embodiment, the remote device is a vehicle.

[0077] According to an embodiment, the set of data is defined by a set of characteristics of the data.

[0078] According to an embodiment, the characteristic includes the model of the remote device.

[0079] According to an embodiment, the characteristic includes a geographical area containing the remote device.

[0080] According to an embodiment, the characteristic includes the use of the characteristics of the remote device.

[0081] According to an embodiment, the instructions further include instructions for: selecting, according to a mapping that associates the classification with the characteristic, a plurality of classifications of data transmitted within the remote device to be included in the set of data.

[0082] According to an embodiment, the instructions further include instructions for: determining a classification name of at least one of the classifications, the classification name being specific to the model of the remote device; and including the classification name in the request for the remote device.

[0083] According to an embodiment, the characteristic is a first characteristic; the classification is a first classification; the mapping associates a second classification with a second characteristic; and the instructions further include instructions for: outputting a prompt to the user to suggest that the user add the second characteristic to the set of characteristics in response to an overlap between the first classification and the second classification.

[0084] According to an embodiment, the characteristic is a first characteristic; and the instructions further include instructions for: outputting a prompt to the user to suggest that the user add a second characteristic to the set of characteristics in response to a correlation between a request for the first characteristic and a request for the second characteristic.

[0085] According to an embodiment, the characteristic includes an event that affects the remote device.

[0086] According to an embodiment, the input is a second input; and the instructions further include instructions for receiving a first input specifying the characteristic.

[0087] According to an embodiment, the instructions further include instructions for: determining a group of the remote devices based on the characteristic; and randomly selecting an actual sample from the group.

[0088] According to an embodiment, the characteristic is a first characteristic; the actual sample is a first actual sample; and the instructions further include instructions for: selecting a second actual sample of the remote device based on a second characteristic related to the first characteristic.

[0089] According to an embodiment, the instructions further include instructions for: receiving an input specifying a maximum value of the cost rating; and determining a candidate confidence value according to a relationship between the confidence level and the cost rating.

[0090] According to an embodiment, the instructions further include instructions for performing the following operations: receiving an input specifying a minimum value of the confidence level; and determining a candidate cost value based on a relationship between the confidence level and the cost rating.

[0091] According to an embodiment, the instructions further include instructions for performing the following operations: selecting the actual samples of the remote devices based on bandwidth limitations of the corresponding remote devices in the actual samples.

[0092] According to an embodiment, the cost rating is a function of the number of remote devices in the actual samples.

[0093] According to an embodiment, the confidence level is a function of the number of remote devices in the actual samples.

[0094] According to the present invention, a method includes: determining a relationship between a confidence level and a cost rating for collecting a set of data from a plurality of remote devices, the confidence level indicating a statistical confidence of the set of data based on candidate samples of the remote devices, the cost rating indicating a cost of collecting the set of data from the candidate samples of the remote devices; then receiving an input selecting a confidence value of the confidence level and a cost value of the cost rating, the confidence value and the cost value being consistent with the relationship between the confidence level and the cost rating; selecting actual samples of the remote devices based on the confidence value; and transmitting a request for the set of data to the remote devices in the actual samples.

Claims

1. A method comprising: determining a relationship between a confidence level for collecting a set of data from a plurality of remote devices and a cost rating, the confidence level indicating a statistical confidence of the set of data based on a candidate sample of the remote devices, and the cost rating indicating a cost of collecting the set of data from the candidate sample of the remote devices; then receiving input selecting a confidence value for the confidence level and a cost value for the cost rating, the confidence value and the cost value being consistent with the relationship between the confidence level and the cost rating; selecting an actual sample of the remote device according to the confidence value; as well as A request for the set of the data is transmitted to the remote device in the actual sample. The method of claim 1 , wherein the remote device is a vehicle. The method of claim 1 , wherein the set of data is defined by a set of characteristics of the data. The method of claim 3 , wherein the characteristic comprises a model number of the remote device.

5. The method of claim 3, wherein the characteristic comprises a geographic area that includes the remote device. The method of claim 3 , wherein the characteristic comprises usage of a feature of the remote device.

7. The method of claim 6, further comprising selecting a plurality of classifications of data transmitted within the remote device for inclusion in the set of data based on a mapping associating the classifications with the characteristics.

8. The method of claim 7, wherein The feature is a first feature; The classification is a first classification; and The mapping associates the second classification with the second feature; The method also includes outputting a prompt to a user in response to an overlap between the first classification and the second classification to suggest that the user add the second feature to the set of characteristics.

9. The method of claim 6, wherein the feature is a first feature; The method also includes outputting a prompt to a user in response to a correlation between the request for the first feature and the request for the second feature to suggest that the user add the second feature to the set of characteristics.

10. The method of claim 3, wherein the characteristic comprises an event affecting the remote device.

11. The method of claim 3, further comprising: determining a population of the remote devices based on the characteristics; as well as The actual sample is randomly selected from the population.

12. The method of claim 1, further comprising: receiving an input specifying a maximum value for the cost rating; as well as A candidate confidence value is determined based on the relationship between the confidence level and the cost rating.

13. The method of claim 1, further comprising: receiving an input specifying a minimum value for the confidence level; as well as A candidate cost value is determined based on the relationship between the confidence level and the cost rating.

14. The method of claim 1, further comprising selecting the actual samples of the remote devices based on bandwidth limitations of corresponding remote devices in the actual samples.

15. A computer comprising a processor and a memory, the memory storing instructions executable by the processor to perform the method of one of claims 1 to 14.