Inference device, inference method, and program
The inference device dynamically selects models based on recognition state patterns to reduce hardware resource consumption and adapt to changing inference scenarios, improving efficiency and adaptability.
Patent Information
- Application Number
- JP2021074174
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-04-26
AI Technical Summary
Conventional CNN models consume significant hardware resources and cannot be dynamically changed during inference, limiting resource efficiency and adaptability.
An inference device that dynamically selects an appropriate model from a plurality of models based on the inference situation, using a selection unit to switch models based on recognition state patterns and state variables.
Reduces hardware resource usage by selecting lightweight models tailored to specific inference situations, enhancing resource efficiency and adaptability.
Smart Images

Figure 0007700501000001 
Figure 0007700501000002 
Figure 0007700501000003
Abstract
Description
Technical Field
[0001] The present invention relates to an inference device, an inference method, and a program.
Background Art
[0002] Conventionally, various tasks such as image recognition have been performed by a machine learning model. Also, when executing a certain task, it has conventionally been performed to select an appropriate machine learning model from among a plurality of machine learning models. For example, Patent Document 1 discloses a technique for selecting a CNN (Convolutional Neural Network) model suitable for a database from among a plurality of CNN models.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, it is common for a CNN model to use a large amount of hardware resources such as RAM (Random Access Memory). For this reason, in Patent Document 1, it is considered that no matter which CNN model is selected from among a plurality of CNN models, it is not possible to significantly reduce the amount of hardware resources used. Also, in Patent Document 1, a CNN model is selected before executing inference, and it is not possible to dynamically change the CNN model during inference.
[0005] On the other hand, depending on the situation during inference, not only a general-purpose machine learning model such as a CNN model but also a lighter machine learning model specialized for a specific situation may be applicable. Therefore, when the situation during inference is a specific situation, it is considered that the amount of hardware resources used can be greatly reduced by performing inference using a lightweight machine learning model specialized for that situation. Also, since the situation during inference can change moment by moment, it is considered preferable that the machine learning model used for inference can be dynamically changed.
[0006] One embodiment of the present invention has been made in view of the above points, and an object thereof is to dynamically select an appropriate model from among a plurality of models according to the situation during inference.
Means for Solving the Problem
[0007] To achieve the above object, an inference device according to one embodiment is an inference device that repeatedly executes inference of a predetermined task for each of a plurality of data, and an inference unit that executes inference of the task on the data using a selected model indicating a model selected from among a plurality of models, and a selection unit that selects a selected model for executing inference in the next repetition from among the plurality of models based on the result of the inference.
Effects of the Invention
[0008] An appropriate model can be dynamically selected from among a plurality of models according to the situation during inference.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0010] Hereinafter, an embodiment of the present invention will be described. In this embodiment, for an image recognition task of recognizing the type and position of an object in an image, an appropriate model is dynamically selected from a plurality of models according to the situation at the time of inference of the task, and then the inference is performed by switching to this selected model. A dynamic model switching inference device 10 will be described. Note that in this embodiment, the model means a machine learning model (image recognition model) or an image recognition method applicable to the image recognition task.
[0011] Here, in this embodiment, each model is configured to recognize the type of object (object type) in the image and its position in the image of the object (position in the image), and calculate the recognition rate thereof. In other words, each model outputs the recognition rate, the object type, and the position in the image. The recognition rate is an index value representing the recognition accuracy of features by the model, and represents, for example, the certainty of the object type and the position in the image output by the model.
[0012] Since the object type and the position in the image are features in the image recognized by the model, hereinafter, the object type and the position in the image will also be referred to as "features". Also, since the recognition rate, the object type, and the position in the image are variables representing the recognition state of the model, the recognition rate, the object type, and the position in the image will also be referred to as "state variables", respectively.
[0013] However, in addition to the recognition rate, object type, and position in the image, various variables representing the recognition state of the model in the image recognition task may also be used as state variables. For example, variables representing the distance from the imaging device to the object, conditions that affect the recognition rate of image recognition (such as lighting conditions in the shooting environment, time of day, etc.) may be used as state variables.
[0014] <Hardware Configuration of Dynamic Model Switching Inference Device 10> First, the hardware configuration of the dynamic model switching inference device 10 according to the present embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing an example of the hardware configuration of the dynamic model switching inference device 10 according to the present embodiment.
[0015] As shown in FIG. 1, the dynamic model switching inference device 10 according to the present embodiment is realized by a general computer or computer system, and includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a processor 105, and a memory device 106. These hardware components are communicably connected to each other via a bus 107.
[0016] The input device 101 is, for example, a keyboard, a mouse, a touch panel, or the like. The display device 102 is, for example, a display or the like. Note that the dynamic model switching inference device 10 may not have at least one of the input device 101 and the display device 102.
[0017] The external I / F 103 is an interface with an external device such as a recording medium 103a. The dynamic model switching inference device 10 can read from and write to the recording medium 103a via the external I / F 103. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), a USB (Universal Serial Bus) memory card, and the like.
[0018] The communication I / F 104 is an interface for connecting the dynamic model switching inference device 10 to a communication network. The processor 105 is various arithmetic units such as, for example, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The memory device 106 is various storage devices such as, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a RAM, a ROM (Read Only Memory), a flash memory, etc.
[0019] By having the hardware configuration shown in FIG. 1, the dynamic model switching inference device 10 according to the present embodiment can realize the dynamic model switching inference process described later. Note that the hardware configuration shown in FIG. 1 is an example, and the dynamic model switching inference device 10 may have other hardware configurations. For example, the dynamic model switching inference device 10 may have a plurality of processors 105 or may have a plurality of memory devices 106.
[0020] Also, the dynamic model switching inference device 10 according to the present embodiment may be, for example, an embedded device or the like that has restrictions on hardware resources compared to a general computer or computer system.
[0021] <Functional Configuration of the Dynamic Model Switching Inference Device 10> Next, the functional configuration of the dynamic model switching inference device 10 according to the present embodiment will be described with reference to FIG. 2. FIG. 2 is a diagram showing an example of the functional configuration of the dynamic model switching inference device 10 according to the present embodiment.
[0022] As shown in FIG. 2, the dynamic model switching inference device 10 according to the present embodiment includes an inference unit 201, a state variable acquisition unit 202, a recognition state pattern determination unit 203, and a model selection unit 204. Each of these units is realized, for example, by a process in which one or more programs installed in the dynamic model switching inference device 10 cause the processor 105 to execute.
[0023] In addition, the dynamic model switching inference device 10 according to this embodiment has a storage unit 205. The storage unit 205 is realized by, for example, a memory device 106. Note that the storage unit 205 may be realized by, for example, a database server or the like connected to the dynamic model switching inference device 10 via a communication network.
[0024] The inference unit 201 takes the image data as input and performs inference processing using the model selected by the model selection unit 204 (hereinafter also referred to as the selected model), recognizes the object type and the position in the image represented by the image data, and calculates the recognition rate.
[0025] The state variable acquisition unit 202 acquires a state variable value representing the inference result (that is, the recognition rate, object type, and position in the image) by the inference unit 201.
[0026] The recognition state pattern determination unit 203 refers to the recognition state pattern information stored in the storage unit 205 and determines the current recognition state pattern from the state variable value (and its change) acquired by the state variable acquisition unit 202. Here, the recognition state pattern is a pattern representing the situation at the time of inference. As will be described later, for example, it is defined by a combination of high / low recognition rate, large / small change in object type, and large / small change in position in the image. Details of the recognition state pattern information will be described later.
[0027] The model selection unit 204 refers to the model type information stored in the storage unit 205 and determines the selected model from the recognition state pattern determined by the recognition state pattern determination unit 203. Here, the model type information is information in which an appropriate model is defined for each recognition state pattern. An appropriate model is, for example, a model that can perform image recognition with high accuracy and in a lightweight manner in the corresponding recognition state pattern (that is, a lightweight model specialized for the recognition state pattern). Details of the model type information will be described later.
[0028] The memory unit 205 stores various types of information (e.g., recognition state pattern information, model type information, model data of each model, etc.). In addition to the recognition state pattern information, model type information, and model data, the memory unit 205 also stores various information necessary for executing the dynamic model switching inference process described later.
[0029] ≪Recognition State Pattern Information≫ Next, an example of the recognition state pattern information stored in the memory unit 205 will be described with reference to FIG. 3. FIG. 3 is a diagram showing an example of the recognition state pattern information.
[0030] In the recognition state pattern information shown in FIG. 3, five recognition state patterns are defined according to the combination of the recognition rate, the change in the object type, and the change in the position in the image. Here, the change in the object type refers to the change in the value of the object type between the current time and the previous time. Similarly, the change in the position in the image refers to the change in the value of the position in the image between the current time and the previous time. In FIG. 3, "×" indicates a low recognition rate, a large change in the object type, and a large change in the position in the image, and "〇" indicates a high recognition rate, a small change in the object type, and a large position in the image. Also, "-" means not defined.
[0031] Recognition state pattern 1 is a recognition state pattern representing the case where the recognition rate is "×". In this case, the change in the object type and the change in the position in the image are not defined. This is because when the recognition rate is low, it is highly likely that the object type and the position in the image are incorrect.
[0032] The recognition state pattern 2 represents the case where the recognition rate is "〇", the change in the object type is "×", and the change in the position in the image is "×". The recognition state pattern 3 represents the case where the recognition rate is "〇", the change in the object type is "×", and the change in the position in the image is "〇". The recognition state pattern 4 represents the case where the recognition rate is "〇", the change in the object type is "〇", and the change in the position in the image is "×". The recognition state pattern 5 represents the case where the recognition rate is "〇", the change in the object type is "〇", and the change in the position in the image is "〇".
[0033] In the example shown in Fig. 3, when the recognition rate is "×", only the recognition state pattern 1 is used, but this recognition state pattern 1 may be further subdivided. That is, it may be subdivided into recognition state pattern 1-1 where the recognition rate is "×", the change in the object type is "×", and the change in the position in the image is "×", recognition state pattern 1-2 where the recognition rate is "×", the change in the object type is "〇", and the change in the position in the image is "×", recognition state pattern 1-3 where the recognition rate is "×", the change in the object type is "〇", and the change in the position in the image is "〇", and recognition state pattern 1-4 where the recognition rate is "×", the change in the object type is "〇", and the change in the position in the image is "〇".
[0034] Also, in the example shown in Fig. 3, the combination of high / low recognition rate, large / small change in the object type, and large / small change in the position in the image is used, but it may be further subdivided. That is, for example, it may be subdivided into three levels: high / medium / low recognition rate. Similarly, at least one of the change in the object type and the change in the position in the image may be subdivided into three levels: large / medium / small. However, the subdivision into three levels is just an example, and it may be further subdivided more finely (that is, subdivided into four or more levels).
[0035] Also, the example shown in Fig. 3 is an example of the recognition state pattern when the three of the recognition rate, the object type, and the position in the image are state variables, and if the number of state variables increases, the number of recognition state patterns may also increase.
[0036] ≪Model type information≫ Next, an example of the model type information stored in the memory unit 205 will be described with reference to FIG. 4. FIG. 4 is a diagram showing an example of the model type information.
[0037] In the model type information shown in FIG. 4, appropriate models A to E are defined for each of recognition state patterns 1 to 5.
[0038] Recognition state pattern 1 is a recognition state pattern with a low recognition rate. Therefore, for recognition state pattern 1, a general-purpose model A that can recognize a plurality of features (object type and position in the image) with high accuracy is used as an appropriate model. Here, as model A, for example, there are a model that automatically extracts feature amounts and a model that manually selects or chooses the feature amounts of interest and combines a plurality of feature amounts. Examples of the former model include a CNN model and the like. On the other hand, examples of the latter model include a model that combines a plurality of feature amounts such as SIFT (Scale-Invariant Feature Transform) feature amounts and HOG (Histgram Of Gradient) feature amounts. Note that since a general-purpose model A such as a CNN model is used in recognition state pattern 1, the effect of reducing the usage amount of hardware resources cannot be expected. This is because in recognition state pattern 1, improvement of recognition performance should be prioritized over reduction of the usage amount of hardware resources.
[0039] Recognition state patterns 2 to 4 are recognition state patterns with a high recognition rate while having a large change in specific features.
[0040] Since Recognition State Pattern 2 has significant changes in both the object type and the position in the image, Model B, which can recognize these two features in a lightweight manner, is used as the appropriate model. Examples of Model B include an image recognition model or method using SIFT features, an image recognition model or method using SURF (Speeded-Up Robust Features) features, etc. However, since Recognition State Pattern 2 has significant changes in both the object type and the position in the image, a CNN model may be used as Model B, or an image recognition model with recognition performance equivalent to that of a CNN model may be used as Model B.
[0041] Since Recognition State Pattern 3 has a significant change in the object type, Model C, which can recognize the object type in a lightweight manner, is used as the appropriate model. Examples of Model C include an image recognition model or method using SIFT features, an image recognition model or method using HOG features, etc.
[0042] It is known that an image recognition model or method using SIFT features is excellent in grasping fine features, while an image recognition model or method using HOG features is excellent in grasping general features. Therefore, for example, when recognizing the type of an object of the same item (e.g., when recognizing the type of pattern on a cup), an image recognition model or method using SIFT features is excellent. On the other hand, for example, when recognizing the type of an object item (e.g., when recognizing whether an item is a cup or not), an image recognition model or method using HOG features is excellent.
[0043] Since Recognition State Pattern 4 has a significant change in the position in the image, Model D, which can recognize the position in the image in a lightweight manner, is used as the appropriate model. Here, a model that can be recognized in a lightweight manner (hereinafter also referred to as a lightweight model) refers to a model that is specialized in recognizing specific features and can perform image recognition with resource savings. Examples of Model D include an image recognition model or method using SURF features, an image recognition model or method using HOG features, etc.
[0044] The recognition state pattern 5 is a recognition state pattern with a high recognition rate and small changes in each feature. In this case, the lightest model E is used as the appropriate model. Examples of model E include an image recognition model or an image recognition method using SURF feature amounts. However, other lightweight image recognition models or image recognition methods may also be used as model E.
[0045] Note that it is not necessary for all of models A to E to be different models, and some of the models may be the same model. For example, both model D and model E may be an image recognition model or an image recognition method using SURF feature amounts.
[0046] Also, the specific examples of the above models are just examples, and various image recognition models or image recognition methods other than the image recognition models or image recognition methods using the SIFT feature amounts, SURF feature amounts, or HOG feature amounts described above can be used. For example, image recognition methods such as Haar-like feature amounts and template matching may be used.
[0047] <Dynamic model switching inference process> Next, the dynamic model switching process according to the present embodiment will be described with reference to FIG. 5. FIG. 5 is a flowchart showing an example of the flow of the dynamic model switching inference process. In the following, an index representing time is t, and the dynamic model switching process at a certain time t will be described.
[0048] First, the inference unit 201 acquires the image data at time t (step S101). Note that the inference unit 201 may acquire the image data from another device or equipment (for example, an imaging device etc.) connected to the dynamic model switching inference device 10, or may acquire the image data from the storage unit 205.
[0049] Next, the inference unit 201 uses the image data acquired in step S101 above as input to perform inference processing using the current selection model, and outputs state variable values representing the recognition rate, object type, and position in the image (step S102). These state variable values are stored in the storage unit 205 as the inference result at time t. When t = 1 (i.e., when performing inference processing for the first time), a pre-determined model (for example, a general-purpose model such as a CNN model) may be used as the selection model.
[0050] Next, the state variable acquisition unit 202 acquires the state variable value output in step S102 above (i.e., the state variable value at time t) and the state variable value at the previous time (i.e., the state variable value at time t - 1) (step S103).
[0051] Next, the dynamic model switching inference device 10 executes model selection processing (step S104). In this model selection processing, based on the state variable value at time t or both the state variable value at time t and the state variable value at time t - 1, the selection model to be used in the inference processing at the next time t + 1 is determined. Details of the model selection processing will be described later.
[0052] Next, the inference unit 201 determines whether there is the next image data (step S105). If there is the next image data, after updating time t to t + 1 (step S106), the inference unit 201 returns to step S101 above. As a result, steps S101 to S104 above are repeatedly executed for each time t. On the other hand, if there is no next image data, the inference unit 201 ends the dynamic model switching inference processing.
[0053] ≪Model Selection Processing≫ Next, the model selection processing in step S104 above will be described with reference to FIG. 6. FIG. 6 is a flowchart showing an example of the flow of the model selection processing.
[0054] First, the recognition state pattern determination unit 203 determines whether the recognition rate included in the state variable value at time t is equal to or greater than a predetermined threshold th1 (step S201). When the recognition rate is equal to or greater than the threshold th1, it is determined that the recognition rate is high, while when it is less than the threshold th1, it is determined that the recognition rate is low. However, for example, thresholds th 11 and th 12 are prepared, and when the recognition rate is equal to or greater than th 11 , it is determined that the recognition rate is high, and when it is less than th 12 , it is determined that the recognition rate is low. When it is equal to or greater than th 12 and less than th 11 , it may be the same as the determination result of this step at time t - 1.
[0055] Note that the threshold th1 (or the thresholds th 11 and th 12 ) may be a preset fixed value, or may be a variable value according to the past recognition rate. For example, after calculating the average value μ1 and variance σ1 of the recognition rates at each time from time t - ΔT (where ΔT is a predetermined time width) to time t, th 11 = μ1 + σ1, th 12 = μ1 - σ1 may be used. However, in order to prevent th 11 from becoming too low, for example, with a preset tuning value of R' (e.g., R' = 70% etc.), th 11 = max(μ1 + σ1, R') may be used.
[0056] When it is determined in step S201 above that the recognition rate is low, the recognition state pattern determination unit 203 refers to the recognition state pattern information shown in FIG. 3 and determines that the current recognition state pattern is "recognition state pattern 1" (step S202).
[0057] Then, the model selection unit 204 refers to the model type information shown in FIG. 4 and selects model A corresponding to the recognition state pattern 1 as the selected model (step S203).
[0058] On the other hand, when it is determined in step S201 above that the recognition rate is high, the recognition state pattern determination unit 203 calculates the change in the state variable value (excluding the recognition rate) at time t with respect to time t-1 (step S204). The reason for excluding the recognition rate is that the high / low recognition rate has already been determined in step S201 above.
[0059] Here, an example of the method for calculating the change in the state variable value (excluding the recognition rate) will be described. For example, as shown in the upper diagram of FIG. 7, assume that the state variable value representing the result of performing inference processing by a certain selection model with the image data at time t-1 as input is a recognition rate of "80%", an object type of "0: cup", and an image position of "X = 100, Y = 150". Also, for example, as shown in the lower diagram of FIG. 7, assume that the state variable value representing the result of performing inference processing by a certain selection model with the image data at time t as input is a recognition rate of "85%", an object type of "0: cup", and an image position of "X = 150, Y = 100". Note that the recognition rate can take a value between 0 and 100, the object type can take a category value representing the type of the object, and the image position can take XY coordinate values with the upper left of the image represented by the image data as the origin, the right direction as the positive direction of the X axis, and the downward direction as the positive direction of the Y axis.
[0060] At this time, for example, when the object type at time t is the same as the object type at time t-1, the recognition state pattern determination unit 203 sets the change to "0", and when they are different, it sets the change to "1". Also, for example, the recognition state pattern determination unit 203 calculates the change in the image position by subtracting the X coordinate value and the Y coordinate value of the image position at time t-1 from the X coordinate value and the Y coordinate value of the image position at time t, respectively. As a result, the change in the object type and the change in the image position are as shown in FIG. 8. That is, the change in the object type at time t with respect to time t-1 is "0", and the change in the image position at time t with respect to time t-1 is "X = 50, Y = 0".
[0061] Next, the recognition state pattern determination unit 203 performs threshold determination on the change in the state variable value other than the recognition rate, and determines whether the change is large or small (step S205). That is, the recognition state pattern determination unit 203 determines whether the change in the object type is large or small, and whether the change in the position in the image is large or small, respectively. Hereinafter, the details of the case where each of the change in the object type and the change in the position in the image is determined to be large or small will be described.
[0062] · When determining whether the change in the object type is large / small For example, after calculating the change rate Rv of the object type, the recognition state pattern determination unit 203 determines that the change in the object type is large when the change rate Rv is equal to or greater than a predetermined threshold th2, and determines that the change in the object type is small when it is less than the threshold th2. Here, the change rate Rv of the object type is calculated as Rv = (n / N) × 100, where n is the number of changes in the object type from time t - ΔT to time t, and N is the number of time indices included in the time width ΔT.
[0063] However, for example, threshold th 21 and threshold th 22 are prepared, and when the change rate Rv of the object type is equal to or greater than threshold th 21 , it is determined that the change in the object type is large, and when it is less than threshold th 22 , it is determined that the change in the object type is small. When it is less than threshold th 21 and equal to or greater than threshold th 22 , it may be the same as the determination result of this step at time t - 1 (or, if this step was not executed at time t - 1, the latest determination result of this step at a previous time).
[0064] Note that the threshold th2 (or thresholds th 21 and th 22 ) may be a preset fixed value, or may be a variable value according to the past change rate Rv. For example, after calculating the average value μ2 and variance σ2 of the change rate Rv of the object type at each time from time t - ΔT to time t, th 21= μ2 + σ2, th 22 It may be μ2 - σ2. However, th 21 To prevent it from becoming too low, for example, with a predetermined tuning value as Rv' (e.g., Rv' = 30% etc.), th 21 It may be th = max(μ2 + σ2, Rv').
[0065] · When determining whether the change in the position in the image is large / small For example, the recognition state pattern determination unit 203 calculates the change rate Rd of the position in the image. When this change rate Rd is equal to or greater than a predetermined threshold th3, it determines that the change in the position in the image is large. When it is less than the threshold th3, it determines that the change in the position in the image is small. Here, the change rate Rd of the position in the image is, for example, the higher one of the change rate Rdx in the X-axis direction and the change rate Rdy in the Y-axis direction between time t - ΔT and time t (that is, Rd = max(Rdx, Rdy)). Also, Rdx is calculated as Rdx = (dx / Px) × 100, where dx is the movement amount in the X-axis direction between time t - ΔT and time t, and Px is the total number of pixels in the X-axis direction of the image represented by the image data. Similarly, Rdy is calculated as Pdy = (dy / Py) × 100, where dy is the movement amount in the Y-axis direction between time t - ΔT and time t, and Py is the total number of pixels in the Y-axis direction of the image represented by the image data.
[0066] However, for example, threshold th 31 and threshold th 32 are prepared. When the change rate Rd of the position in the image is equal to or greater than the threshold th 31 it is determined that the change in the position in the image is large. When it is less than the threshold th 32 it is determined that the change in the position in the image is small. When it is less than the threshold th 31 and greater than or equal to the threshold th 32 it may be the same as the determination result of this step at time t - 1 (or, if this step was not executed at time t - 1, the latest determination result of this step at a previous time).
[0067] Note that the threshold th3 (or the threshold th31 and threshold th 32 ) may be a preset fixed value, or may be a variable value according to the past change rate Rd. For example, after calculating the average value μ3 and variance σ3 of the change rate Rd of the position in the image at each time from time t - ΔT to time t, th 31 = μ3 + σ3, th 32 = μ3 - σ3 may be used. However, in order to prevent th 31 from becoming too low, for example, using a preset tuning value as Rd' (for example, Rd' = 200% etc.), th 31 = max(μ3 + σ3, Rd') may be used.
[0068] Next, the recognition state pattern determination unit 203 refers to the recognition state pattern information shown in FIG. 3 and determines which of the recognition state patterns 1 to 5 the current recognition state pattern is (step S206). That is, when the change in the object type is large and the change in the position in the image is also large, the recognition state pattern determination unit 203 determines "recognition state pattern 2", when the change in the object type is large and the change in the position in the image is small, it determines "recognition state pattern 3", when the change in the object type is small and the change in the position in the image is large, it determines "recognition state pattern 4", and when the change in the object type is small and the change in the position in the image is also small, it determines "recognition state pattern 5".
[0069] Next, the model selection unit 204 refers to the model type information shown in FIG. 4 and selects the model corresponding to the recognition state pattern determined in step S206 above as the selected model (step S207). That is, when it is determined as "recognition state pattern 2" in step S206 above, the model selection unit 204 selects model B, when it is determined as "recognition state pattern 3", it selects model C, when it is determined as "recognition state pattern 4", it selects model D, and when it is determined as "recognition state pattern 5", it selects model E as the selected model.
[0070] Subsequently, following the above step S203 or step S207, the model selection unit 204 reads the model data of the selected model from the storage unit 205 (step S208). As a result, the selected model will be used in the inference process at the next time t + 1.
[0071] <Summary> As described above, the dynamic model switching inference device 10 according to the present embodiment can dynamically switch the selected model to a model corresponding to the pattern in response to the result of the inference process by the current selected model and the recognition state pattern representing the pattern of its change, targeting an image recognition task. Thereby, for example, when the recognition rate is low, a general-purpose model with a higher recognition rate can be selected as the selected model, or when the change in a certain feature is large, a lightweight model specialized for the recognition of that feature can be selected as the selected model. Therefore, particularly in a situation where a certain recognition rate or higher is obtained, since the inference process can be performed by a lightweight model specialized for the recognition of a specific feature, it is possible to reduce the amount of hardware resources used compared to the case where the inference process is always performed by a general-purpose model.
[0072] In the present embodiment, the image recognition task is targeted, but this is just an example, and it is applicable to various tasks that repeatedly execute inference processing by a model using data as input. For example, it is applicable to various tasks such as data classification, regression, prediction, anomaly detection, etc.
[0073] The present invention is not limited to the specifically disclosed above embodiments, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the description of the claims.
Explanation of Reference Numerals
[0074] 10 Dynamic model switching inference device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 Processor 106 Memory device 107 Bus 201 Inference unit 202 State variable acquisition unit 203 Recognition state pattern determination unit 204 Model selection unit 205 Storage unit
Claims
1. An inference device that repeatedly executes inference of a predetermined task for each of a plurality of data, comprising: an inference unit that executes inference of the task on the data using a selected model indicating a model selected from among a plurality of models; a selection unit that selects, from among the plurality of models, a selected model for executing inference in the next iteration based on the result of the inference; wherein the result of the inference includes a recognition result for each of one or more features of the data and a recognition rate representing the certainty of the recognition result; the selection unit selects, as the selected model for executing inference in the next iteration, a model corresponding to the pattern according to the pattern of the recognition result and the recognition rate included in the result of the inference in the current iteration.
2. The selection unit selects, as the model corresponding to the pattern, a general-purpose model capable of recognizing the one or more features as the selected model when the recognition rate is less than a predetermined threshold value. The inference device according to claim 1.
3. The selection unit selects, as the model corresponding to the pattern, a lightweight model capable of recognizing a specific feature as the selected model when the recognition rate is greater than or equal to the threshold value. The inference device according to claim 1.
4. The pattern is represented by a combination of whether the recognition rate is greater than or equal to the threshold value, whether there is a change between the recognition result included in the result of the inference in the current iteration and the recognition result included in the result of the inference in the previous iteration, or the magnitude of the change amount. The inference device according to any one of claims 1 to 3.
5. The task is an image recognition task, the one or more features include at least one of the type of an object in the image represented by the data and the position of the object in the image, the pattern is represented by a combination of whether the recognition rate is greater than or equal to the threshold value, whether there is a change between the type of the object included in the result of the inference in the current iteration and the type of the object included in the result of the inference in the previous iteration, and the magnitude of the change amount between the position of the object included in the result of the inference in the current iteration and the position of the object included in the result of the inference in the previous iteration. The inference device according to claim 4.
6. An inference method for repeatedly executing inference of a predetermined task for each of a plurality of data, comprising: An inference procedure for performing inference of the task on the data by a selected model indicating a model selected from among a plurality of models; A selection procedure for selecting, from among the plurality of models, a selected model for performing inference in a next iteration based on the result of the inference; The computer executes; The result of the inference includes a recognition result for each of one or more features of the data and a recognition rate representing the certainty of the recognition result; The selection procedure is; An inference method of selecting, as a selected model for performing inference in the next iteration, a model corresponding to the pattern according to the pattern of the recognition result and the recognition rate included in the result of the inference in the current iteration.
7. A program for repeatedly performing inference of a predetermined task on each of a plurality of data, An inference procedure for performing inference of the task on the data by a selected model indicating a model selected from among a plurality of models; A selection procedure for selecting, from among the plurality of models, a selected model for performing inference in a next iteration based on the result of the inference; Causing the computer to execute; The result of the inference includes a recognition result for each of one or more features of the data and a recognition rate representing the certainty of the recognition result; The selection procedure is; A program that selects, as a selected model for performing inference in the next iteration, a model corresponding to the pattern according to the pattern of the recognition result and the recognition rate included in the result of the inference in the current iteration.
Citation Information
Patent Citations
Determination device and determination method for convolutional neural network model for database
JP2018092614A
Behavior recognition method, behavior recognition program, and behavior recognition device
JP2020071665A
Recognition device, recognition method and program
JP2020160814A
Information processing device for vehicle
JP2021018593A