Visual and tactile sensor device, texture evaluation method using the visual and tactile sensor device, and texture evaluation program
The visual-tactile sensor device addresses the lack of texture evaluation in existing tactile sensors by using a marker-based system to analyze object characteristics, achieving precise texture assessment of diverse materials.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SAN EI GEN F F I INC
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing tactile sensors, such as those described in Patent Document 2, are not optimized for texture evaluation and lack the capability to effectively assess the shape, size, distribution, hardness, and texture of objects, including rigid and deformable bodies like food.
A visual-tactile sensor device with a partially light-transmitting visor-tactile receiving portion containing markers, a motion driving mechanism, and an imaging unit that captures marker positions during predetermined operations, coupled with an evaluation unit to analyze marker displacement for texture evaluation.
Enables qualitative and quantitative assessment of shape, size, distribution, hardness, and texture of various objects, including rigid and deformable bodies, with high accuracy and discrimination of surface textures.
Smart Images

Figure 2026076522000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a visual tactile sensor device, a texture evaluation method using the visual tactile sensor device, and a texture evaluation program.
Background Art
[0002] Patent Document 1 discloses a machine learning system for learning a texture evaluation model for evaluating the texture of food. The machine learning system includes a pressing device for pressing a sample, an image acquisition unit that acquires pressure distribution images of a plurality of frames showing a time-series pressure distribution received from the sample when the sample is pressed, and generates one input image by connecting the pressure distribution images of the plurality of frames, and a learning unit that learns the texture evaluation model by a convolutional neural network based on the input image.
[0003] Patent Document 2 discloses a tactile sensor device. The tactile sensor device includes a compound eye imaging device in which a plurality of compound eye imaging elements are two-dimensionally arranged on a flexible sheet, a lighting device that illuminates the imaging region of the compound eye imaging device, and a deformation layer in which a marker is formed in the imaging region.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
[0007] The primary objective of this disclosure is to provide a visual-tactile sensor device capable of evaluating texture. The second objective is to provide a texture evaluation method and a texture evaluation program using a visual and tactile sensor device. [Means for solving the problem]
[0008] A visual and tactile sensor device (A1) suitable for texture evaluation is: A visor-tactile receiving portion (1) that is partially or entirely light-transmitting in the thickness direction, A plurality of markers (M_i, i=1, ...n) are arranged at predetermined positions inside one surface (the side that contacts the object to be measured) in the thickness direction of the visual-tactile receiving portion (1), A motion driving means (30) is provided to bring the object to be measured into contact with the optic and tactile receiving part (1) on the marker (M_i) side and perform a predetermined operation. While the motion driving means (30) is causing the object to be measured to perform the predetermined action, the imaging unit (40) captures the positional state of the marker (M_i) from the side of the visual-tactile receiving unit (1) that is not in contact with the object to be measured, An evaluation unit (50) analyzes the image data captured by the imaging unit (40) and evaluates one or more of the shape (surface characteristics), size, distribution, hardness, and texture of the object to be measured, It may also be equipped with.
[0009] The "object of measurement" is not particularly limited and may include, for example, a rigid body, a deformable body, or food.
[0010] The evaluation unit (50) is, The shape (surface characteristics), size, distribution, hardness, and texture of the object to be measured may be evaluated qualitatively or quantitatively. For example, the quantitative evaluation may involve evaluating estimated values obtained by quantifying the levels in stages.
[0011] The optic-tactile receptor (1) and the marker (M_i) function like human tactile receptors. The predetermined operation may be one or more of the following operations combined: (a) a pressing operation that pushes the marker (M_i) in to the extent that it is displaced from its initial position; (b) a pressing and rotating operation that rotates the marker in the pressed position (fixed pressing position); (c) a pressing and translation operation that translates the marker in a predetermined direction in the pressed position (fixed pressing position); (d) an operation that rotates the marker while pressing it; (e) an operation that translates the marker while pressing it.
[0012] The optic and tactile receiving portion (1) may have a uniform hardness or elastic modulus, or it may be a combination of multiple members with different hardnesses and elastic moduli. The area where the marker is placed on the visual-tactile receiving part (1) may be made of a material with a hardness similar to that of a human tongue, such as transparent silicone rubber. A hardness similar to that of a human tongue may have an elastic modulus of 10kPa to 200kPa at a compressive strain of 20%. The compressive stress is determined in accordance with the compression test method of JIS K 6254.
[0013] The markers (M_i) may be arranged, for example, in a grid, radial, or concentric multi-layered circular pattern. The adjacent markers (M_i) may be arranged, for example, at equal or unequal intervals. The plurality of protrusions (P_i) may be arranged, for example, in a grid, radial, or concentric multi-layered circular pattern. The adjacent plurality of protrusions (P_i) may be arranged, for example, at equal or unequal intervals.
[0014] The marker (M_i) may be, for example, spherical in shape with a diameter of 0.4 mm to 2 mm. The protrusion (P_i) may be, for example, hemispherical, quadrangular pyramidal, or conical with a radius of 1 mm to 3 mm.
[0015] The evaluation unit (50) includes a marker displacement acquisition unit (51) that obtains marker displacement information for each of the markers (M_i) from the acquired time-series image data, and a first estimation unit (512) that evaluates one or more of the shape (surface characteristics), size, distribution, firmness, and texture of the measurement target based on the marker displacement information obtained by the marker displacement acquisition unit (51). The evaluation unit (50) further includes a feature amount calculation unit (511) that calculates a feature amount of the marker displacement obtained by the marker displacement acquisition unit (51), and the first estimation unit (512) may evaluate one or more of the shape (surface characteristics), size, distribution, firmness, and texture of the measurement target based on the marker displacement information and / or the feature amount of the marker displacement. Examples of the "marker displacement information" include, for example, the contact area (A) between the measurement target and the visual and tactile reception unit (1), the marker displacement amount (including the average value), the change rate (displacement speed), the marker displacement vector, the quantized heat map, and the like. Examples of the "feature amount of the marker displacement" include, for example, the average frequency (f_ave), the average amplitude (A_ave), and the like. The "texture" may be one or more of five types: "chewy feeling", "slippery feeling", "sticky feeling", "rough feeling", and "residual feeling".
[0016] The evaluation unit (50) Using a measurement object (sample) for which one or more of shape (surface characteristics), size, distribution, hardness, and texture are known, one or more types of data selected from marker displacement information and feature amounts of marker displacement obtained from image data captured by the imaging unit (40) of the visual tactile sensor device (A1), and one or more types of data of the shape (surface characteristics), size, distribution, hardness, and texture as teacher data (also referred to as "training data"), a storage device (53) that stores a learning model (52) (also referred to as a "classifier") generated by an intelligent information processing technique, A second estimation unit (54) that uses, as input data, one or more types of data selected from marker displacement information and feature amounts of marker displacement obtained from image data captured by the imaging unit (40), and uses the learning model (52) to identify and classify one or more of the shape (surface characteristics), size, distribution, hardness, and texture may be provided.
[0017] The visual tactile sensor device (A1) may have a learning execution unit that inputs the teacher data into the learning model (52) and executes learning. The learning execution unit may be configured to perform re-learning using different teacher data.
[0018] The visual tactile sensor device (A1) may further have an accuracy evaluation unit that calculates accuracy results of accuracy rate ("correct answer rate" also), precision rate, and / or recall rate from the estimation results obtained from the learning model (52).
[0019] The visual tactile sensor device (A1) may have an output unit (60) that outputs the results by the first estimation unit (512) and / or the second estimation unit (54). The output may be composed of various means such as displaying on a display device, transmitting to another device, printing, emitting sound, storing in a storage device, etc.
[0020] The visual tactile reception unit (1) may have a structure of two or more layers in the thickness direction. The portion of the visual-tactile receiving part (1) where the marker (M_i) is placed is made of at least a soft material. When the aforementioned visual-tactile receiving portion (1) has a two-layer structure, A first member (10) that is partially or entirely light-transmitting in the thickness direction, The present invention may also have a second member (20) that is softer than the first member (10) and laminated, and is partially or entirely light-transmitting. The plurality of markers (M_i) may be arranged at predetermined positions inside one surface (the side in contact with the object to be measured) in the thickness direction of the second member (20). The motion driving means (30) may bring the object to be measured into contact with the second member (20) on the marker (M_i) side and perform a predetermined operation. The imaging unit (40) may capture the positional state of the marker (M_i) from the first member (10) side while the motion driving means (30) is causing the measurement target to perform the predetermined operation. The second member (20) may have a flat surface on the side that contacts the object to be measured. The second member (20) may have a plurality of protrusions (P_i, i=1, ...n) on one surface (the side that contacts the object to be measured), and one or more markers (M_i) may be provided on each of the protrusions (P_i).
[0021] The texture evaluation method using the visual and tactile sensor device (A1) is as follows: The method may include a first evaluation step in which marker displacement information and / or marker displacement feature quantities are obtained for each marker (M_i) from the acquired time-series image data, and one or more of the shape (surface properties), size, distribution, hardness, and texture of the object to be measured are evaluated based on the marker displacement information and / or marker displacement feature quantities.
[0022] The texture evaluation method using the visual and tactile sensor device (A1) is as follows: The method may include a second evaluation step in which, using a measurement target (sample) for which one or more of the following are known, one or more data selected from marker displacement information and marker displacement feature quantities obtained from image data captured by the imaging unit (40) of a visual-tactile sensor device (A1), and one or more data selected from the shape (surface characteristics), size, distribution, hardness, and texture, are input into a learning model (52) generated by intelligent information processing technology, which is used as training data (also called "training data"), and one or more data selected from marker displacement information and marker displacement feature quantities obtained from image data captured by the imaging unit (40), to identify and classify one or more of the shape (surface characteristics), size, distribution, hardness, and texture.
[0023] Information processing equipment, At least one processor, The processor includes a memory for storing instructions that can be executed by the processor, The processor is an information processing device that realizes each step of the texture evaluation method by executing executable instructions.
[0024] The texture evaluation program is This is a program that implements each step of the texture evaluation method using at least one processor.
[0025] A computer-readable recording medium that stores computer instructions is This is a computer-readable recording medium that stores a learning model (52) generated by intelligent information processing technology using marker displacement information and one or more data selected from the marker displacement features obtained from image data captured by the imaging unit (40) of a visual-tactile sensor device (A1), and one or more data from the shape (surface characteristics), size, distribution, hardness, and texture, as training data, using a measurement target (sample) for which one or more of the following are known: shape (surface characteristics), size, distribution, hardness, and texture.
[0026] Examples of "intelligent information processing technologies" include machine learning, deep learning, reinforcement learning, and deep reinforcement learning. The algorithms used for machine learning, deep learning, reinforcement learning, and deep reinforcement learning are not particularly restricted, and conventional algorithms may be used. For supervised learning, various algorithms such as linear regression, generalized linear models, support vector regression, Gaussian process regression, ensemble methods, decision trees, neural networks, support vector machines, discriminant analysis, Naive Bayes, and nearest neighbor methods may be employed.
[0027] Each element of the device may consist of an information processing device (e.g., a computer, server) having memory, a processor, and software programs, or dedicated circuits, firmware, etc. The information processing device may be on-premises, in the cloud, or a combination of both. [Brief explanation of the drawing]
[0028] [Figure 1] This figure shows an example of a visual-tactile sensor device according to Embodiment 1. [Figure 2] This figure shows an example of a flat, plate-shaped first member, second member, and marker without protrusions. [Figure 3] This figure shows an example of a first member with a protrusion, a second member, and a marker. [Figure 4] This figure shows an example of the results of the pressing operation on the gel-like food in Example 1. [Figure 5] This figure shows an example of the results of the pressing operation on gel-like foods in Examples 2 and 3. [Figure 6A] This shows an example of estimation using training data and a learning model. [Figure 6B] This figure shows an example of a quantized heatmap based on marker displacement. [Figure 7] This figure shows an example of the estimated results for the four textures. [Figure 8A] This figure shows an example of a rigid object. [Figure 8B]This figure shows an example of marker displacement vectors and indentation amount x for sensors without protrusions and sensors with protrusions. [Figure 8C] This figure shows an example of marker displacement, its average frequency f_ave, and average amplitude A_ave. [Figure 9A] This figure shows an example of marker displacement vectors and indentation amount x for sensors without protrusions and sensors with protrusions. [Figure 9B] This figure shows an example of marker displacement, its average frequency f_ave, and average amplitude A_ave. [Figure 10A] This figure shows an example of a rigid object having 42 patterns of hemispherical protrusions. [Figure 10B] This figure shows an example of a learning model. [Figure 10C] The results of verifications 1 and 2 are shown in the figure. [Figure 11] This figure shows the state of the marker displacement vector from the initial state up to 5 reciprocating rotations. [Figure 12] This figure shows the change over time in the sum of marker displacements. [Modes for carrying out the invention]
[0029] (Embodiment 1) Figure 1 shows an example configuration of the visual-tactile sensor device A1 of Embodiment 1. The visual-tactile sensor device A1 comprises a visual-tactile receiving unit 1, a motion driving means 30, an imaging unit 40, an evaluation unit 50, and an output unit 60.
[0030] The visual-tactile receiving portion 1 of Embodiment 1 has a two-layer structure in the thickness direction. The visual-tactile receiving portion 1 has a first member 10 made of a hard material that is partially or entirely light-transmitting in the thickness direction, and a second member 20 made of a soft material that is softer than the first member 10 and laminated on top of the first member 10, and is partially or entirely light-transmitting. In this embodiment, the first member 10 is made of a transparent hard acrylic material. The second member 20 is made of a transparent soft silicone rubber material. The hardness of the second member 20 needs to be softer than the object being measured, and it is preferable that it has hardness and elasticity similar to that of a human tongue. Its modulus of elasticity is, for example, 10kPa to 200kPa. The second member 10 may be a rectangular prism shape with a thickness of 5 to 20 mm and dimensions of 30 to 50 mm in length and 30 to 50 mm in width in a plan view. The second member 20 may be a rectangular prism shape with a thickness of 5 to 20 mm and dimensions of 30 to 50 mm in length and 30 to 50 mm in width in a plan view. A surrounding wall may be provided that defines the side surface around the second member 20. The surrounding wall may be made of a different material and may be configured to assist the restoring force of the second member 20 at the end of a predetermined operation.
[0031] Multiple markers (M_i) are placed on the surface side of the second member 20 that is in contact with the object to be measured. The markers (M_i) may be black-painted glass spheres. i is from 1 to n. The resolution increases in proportion to n. The markers may be arranged in a grid of 10 to 30 rows and 10 to 30 columns at a depth of 1 mm or less from the surface. The spherical markers may have a diameter of, for example, 0.2 to 1.0 mm. The spacing of the grid arrangement may be, for example, 1.5 to 2.5 mm. The markers are embedded in silicone rubber and move in accordance with the behavior of the silicone rubber when it deforms.
[0032] The motion driving means 30 brings the object to be measured (sample) into contact with the second member 20 on the marker (M_i) side and performs a predetermined operation. The predetermined operation may be one or more of the following operations combined: (a) a pressing operation that pushes the marker (M_i) in to the extent that it is displaced from its initial position; (b) a pressing and rotating operation that rotates it while it is pressed in; (c) a pressing and translating operation that translates it in a predetermined direction while it is pressed in; (d) an operation that rotates while pressing; (e) an operation that translates while pressing.
[0033] The motion driving means 30 may include, for example, a pressing unit 31, a vertical actuator 32, a translation actuator 33, a motor 34, a drive structure, a setting unit 35, and an operation control unit 36, as shown in Figure 1. The pressing unit 31 is configured as a flat plate that presses the object to be measured S into the second member 20. The vertical actuator 32 is connected to the pressing unit 31 and is an actuator that can move the pressing unit 31 in the vertical direction and rotates around a vertical axis. The translation actuator 33 is connected to the pressing unit 31 and is an actuator that translates the pressing unit 31. The motor 34 is a driving means that drives the vertical actuator 32 and the translation actuator 33. The setting unit 35 is a means for arbitrarily setting the pressing speed, pressing amount (z), pressing pressure (f), rotation direction, rotation speed, number of rotations, translation direction, translation speed, and number of translations according to the object to be measured S. The various conditions set in the setting unit 35 are sent to the storage device 53. In an alternative embodiment, the functions of the setting unit 35 may be provided in the evaluation unit 50, and various conditions may be sent to the motion control unit 36 of the motion driving means 30. The motion control unit 36 drives the motor 34 and controls the various actuators 32 and 33 according to the various conditions set in the setting unit 35.
[0034] The motion drive means 30 may consist of a multi-axis robot arm equipped with various sensors instead of individual actuators. The multi-axis robot arm can control actions such as pushing, rotation, and translation. The pushing force f and pushing amount z can be determined by the various sensors. The various setting values for the robot arm operation that match the setting values set in the setting unit 35, along with the data of the pushing force f and pushing amount z measured by the various sensors, are sent to the storage device 53.
[0035] The imaging unit 40, while the dynamic drive means 30 is performing a predetermined operation on the object to be measured S, captures the position state of the marker (M_i) as time-series data from the side of the first member 10 that is not in contact with the object to be measured S. This can be a video or a still image. The imaging unit 40 may be a color camera such as a CCD camera or a CMOS camera. The time-series image data captured by the imaging unit 40 is sent to the storage device 53 and stored in the image data storage unit 531.
[0036] The evaluation unit 50 analyzes the image data captured by the imaging unit 40 and evaluates one or more of the following characteristics of the object to be measured: shape (surface characteristics), size, distribution, hardness, and texture. The evaluation unit 50 also includes a storage device 53. The storage device 53 comprises an image data storage unit 531, a condition storage unit 532, an analysis data storage unit 533, and an evaluation data storage unit 534. The evaluation unit 50 receives various condition data set in the setting unit 35, such as indentation speed, indentation amount z, indentation pressure f, rotation direction, rotation speed, number of rotations, translation direction, translation speed, translation distance, and number of translations, and stores them in the condition storage unit 532. The evaluation unit 50 also receives various setting values for robot arm operation, and data on indentation force f and indentation amount z measured by various sensors, and stores them in the condition storage unit 532.
[0037] (Rule-based evaluation) The evaluation unit 50 includes, for example, a marker displacement acquisition unit 51 and a first estimation unit 512. The marker displacement acquisition unit 51 calculates marker displacement information for each marker (M_i) from the acquired time-series image data. The marker displacement information may be the average value M of the marker displacement, the marker displacement vector, the contact area A, and the rate of change of the marker displacement (displacement velocity).
[0038] The marker displacement acquisition unit 51 analyzes the time-series image data captured by the imaging unit 40. The marker displacement acquisition unit 51 detects the marker (M_i) group by blob analysis. These processes are performed on all time-series image data (also called "frames") obtained by imaging. For example, the marker displacement acquisition unit 51 calculates the contact area A corresponding to the indentation amount z based on the time-series image data stored in the image data storage unit 531 and various data stored in the condition storage unit 532. For example, the marker displacement acquisition unit 51 calculates the displacement (x) of each detected marker (M_i) from its initial position (x0, y0). m , y m The system calculates the displacement and generates a vector of length obtained by multiplying this displacement by a constant. m is the same as the number of frames in the captured time series. The vector of length obtained by multiplying by a constant is called the marker displacement vector. For example, as shown in Figure 2, each marker (M_i) that has been displaced over time and the marker displacement vector are sent to the display device of the output unit 60 and displayed. When displaying the marker displacement vector, it is colored in 8 to 16 steps, for example, depending on the amount of displacement. For example, the marker displacement acquisition unit 51 calculates the average value M of the marker displacement within the contact surface. The data calculated by the marker displacement acquisition unit 51 is stored in the analysis data storage unit 533.
[0039] The feature calculation unit 511 calculates the feature quantities of the marker displacement obtained by the marker displacement acquisition unit 51. The feature quantities of the marker displacement may be the average frequency (f_ave) and average amplitude (A_ave) of the marker displacement. The frequency and amplitude may also be determined by frequency analysis from the time-dependent graph of the indentation amount z. Alternatively, the average frequency and amplitude may be determined by repeating the same predetermined operation multiple times.
[0040] The first estimation unit 512 evaluates one or more of the shape (surface characteristics), size, distribution, hardness, and texture of the object to be measured based on the marker displacement information obtained by the marker displacement acquisition unit 51 and / or the marker displacement features calculated by the feature calculation unit 511. For example, the first estimation unit 512 accesses the evaluation database, matches the marker displacement information or marker displacement feature data as extraction conditions, and outputs data on shape (surface characteristics), size, distribution, hardness, and texture. The evaluation database includes data on shape (surface characteristics), size, distribution, hardness, and texture linked to the marker displacement information or marker displacement features, and is pre-set by experimental values. The data processed by the first estimation unit 512 is stored in the evaluation data storage unit 534. The evaluation database includes both qualitative and quantitative evaluations, and both qualitative and quantitative evaluation data are output.
[0041] (Example 1: Push operation) First component 10: Transparent, rigid acrylic, 4mm thick, 40mm x 40mm in size. Second component 20: Made of transparent, soft silicone rubber, with an elastic modulus of 98 kPa, a thickness of 10 mm, and dimensions of 40 mm x 40 mm, surrounded by a 4 mm thick PLA resin component. Markers (M_i): 442 black-painted glass spheres, 0.6 mm in diameter, arranged in 21 rows and 21 columns, spaced 2 mm apart. Image capture conditions: Resolution 1080 x 1080 pixels, frame rate 30 fps Motion drive mechanism: A 6-axis robot arm with a plunger (pressure part) attached to the end of the arm. Motion conditions: Vertical pressing motion with a speed of 2 mm / s and a pressing depth of z = 9 mm. Measurement target: A cylindrical gel-like food made from gellan gum, 20 mm in diameter and 10 mm in height.
[0042] Figure 4(a) shows the contour of the contact surface between the sensor surface and the object S being measured, and displays only the marker displacement vector within it. Figure 4(b) shows the output results for indentation amount z, contact area A, indentation force f, and average marker displacement M. The contact area A increased with increasing indentation amount z. As the indentation force f progressed, there was a significant decrease around t=1.3s, confirming fracture. A decrease was also observed at t=4.5s when the indentation was completed, confirming stress relaxation. The trend of change in the average marker displacement M was consistent with the indentation force f, confirming both fracture and stress relaxation.
[0043] (Example 2: Push operation) The procedure is the same as in Example 1, except that the object of measurement was a cylindrical agar-based gel food with a diameter of 20 mm and a height of 10 mm. Figure 5(a) shows the marker displacement vectors. A non-uniform group of marker displacement vectors was observed after fracture. Stickiness can be evaluated from the data of the average marker displacement M. For example, the stickiness can be evaluated relatively, such as whether the average marker displacement M is stronger or weaker compared to a reference value.
[0044] (Example 3: Push operation) The procedure is the same as in Example 1, except that the object of measurement was a cylindrical carrageenan gel-like food with a diameter of 20 mm and a height of 10 mm. Figure 5(b) shows the marker displacement vectors. After fracture, the marker displacement vectors spread isotropically, and a group of marker displacement vectors with a concentric distribution from the sensor center was observed. Stickiness can be evaluated from the data of the average marker displacement M. For example, the stickiness can be evaluated relatively, such as whether the average marker displacement M is stronger or weaker compared to a reference value.
[0045] (Evaluation using a learning model) Alternatively, or in addition to the rule-based evaluation described above, an evaluation using a learning model is performed. The evaluation unit 50 includes a marker displacement acquisition unit 51, a feature calculation unit 511, and a second estimation unit 54. The storage device 53 stores the learning model 52. The marker displacement acquisition unit 51 has the functions described above.
[0046] The learning model 52 is a learning model generated by intelligent information processing technology using image data captured by the imaging unit 40 of the visual-tactile sensor device A1 using a measurement target (sample) for which one or more of the following are known: shape (surface characteristics), size, distribution, hardness, and texture, and one or more of the shape (surface characteristics), size, distribution, hardness, and texture of the measurement target, as training data. This embodiment uses a learning model based on multiple regression analysis within linear regression. The data output from the learning model includes both qualitative and quantitative evaluation data.
[0047] (Sensory evaluation of texture) As training data for texture, sensory evaluation values for eight types of gel-like foods (A-H) with different textures, created by combining gellan gum, carrageenan, agar, etc., will be prepared. Sensory evaluation values will be obtained by actual human taste testing through chewing. There will be four texture evaluation items: "chewy" (i=1), "sticky" (i=2), "smooth" (i=3), and "gritty" (i=4), with each item having a sensory evaluation value n i The following are defined: "Chewy texture" refers to the degree to which the texture is soft, stretchy, and pushes back against the tongue before breaking, while "sticky texture" refers to the degree to which the texture adheres to the tongue and is difficult to spread after breaking. These are categorized as textures based on mechanical properties. Additionally, "Smooth texture" refers to the degree to which the surface is smooth before breaking, while "rough texture" refers to the degree to which the surface is rough after breaking. These are categorized as textures based on geometric properties. Sensory evaluation is conducted using the Visual Analog Scale method. A sheet is provided for each texture evaluation item i to mark the degree of sensory impression. This sheet has a 100mm long straight line, with the left end marked "I don't feel any 'XX' sensation at all" and the right end marked "I feel a strong 'XX' sensation." For each food item, the evaluator marks the degree of the texture impression felt on the tongue on the straight line. The marked position is measured as an integer value between 0 and 100mm, and the converted value is the sensory evaluation value n for each food item i. i The evaluation will be conducted by eight evaluators, and the average value will be used as training data. An example of training data is shown in Figure 6A(a).
[0048] (Quantization Heatmap) The marker displacement acquisition unit 51 quantizes the marker displacement amount into, for example, 16 steps and converts it into an integer value from 0 to 15. This is called a quantized heatmap. The quantization may also be performed in 8 to 24 steps. Figures 6A(b) and 6B show an example of a quantized heatmap created based on the marker displacement vector. The process from the start of compression until the gel-like food breaks is called the compression phase, and the process from breakage until the end of compression is called the break phase. Data measurement is performed 9 times for each type of gel-like food, and a total of 72 data points are acquired for 8 types of gel-like foods (A to H).
[0049] (Building a learning model) Figure 6A(c) shows an example of the construction of the learning model 52. The feature calculation unit 511 calculates texture features for each frame of the quantized heatmap using the spatial intensity level-dependent method. An intensity co-occurrence matrix is calculated from the image data. The intensity co-occurrence matrix is a matrix that represents the intensity relationships between pixels in the image. Energy, entropy, inertia, correlation, and local uniformity are calculated as five types of features. For each frame, 8 types of intensity co-occurrence matrices × 5 types of features = a total of 40 types of features are calculated. In both the compression phase and the fracture phase, the mean, standard error, maximum value, minimum value, and range are calculated as representative features. From the above, 400 representative features are obtained for a quantized heatmap of one gel-like food. Adding the number of frames in each of the compression and fracture phases to this, a feature vector F = 402 dimensions is obtained. The feature calculation unit 511 further performs principal component analysis to remove high correlations between features and compresses the dimensionality of the feature vector.
[0050] (Derivation of the multiple regression analysis model) Let K be the total number of texture evaluation items, and let n be the sensory evaluation value of texture evaluation item i (i=1,2,...,K). i In this embodiment, K=4. For each texture evaluation item, the principal component vector y is used as the explanatory variable, and the sensory evaluation value n iA multiple regression model is created with the target variable as . The resulting multiple regression model is stored in the memory device 50 as the learning model 52.
[0051] (Using a Video Vision Transformer) A Video Vision Transformer (hereinafter referred to as the "ViViT model") may be constructed and used as a training model. The ViViT model may be constructed, for example, as described in "Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Luci´c and Cordelia Schmid. “Vivit: A video vision transformer”. In: Proceedings of the IEEE / CVF international conference on computer vision. 2021, pp. 6836-6846".
[0052] The second estimation unit 54 takes one or more data selected from marker displacement information and marker displacement features obtained from image data captured by the imaging unit 40 as input data, and uses the learning model 52 to identify and classify one or more of the following: shape (surface characteristics), size, distribution, hardness, and texture. The confidence level of the classified result is also output.
[0053] (Example 4: Estimation results of the learning model) We will evaluate the created learning model.
[0054] First component 10: Transparent, rigid acrylic, 4mm thick, 40mm x 40mm in size. Second component 20: Made of transparent, soft silicone rubber, with an elastic modulus of 98 kPa, a thickness of 10 mm, and dimensions of 40 mm x 40 mm, surrounded by a 4 mm thick PLA resin component. Markers (M_i): 442 black-painted glass spheres, 0.6 mm in diameter, arranged in 21 rows and 21 columns, spaced 2 mm apart. Image capture conditions: Resolution 1080 x 1080 pixels, frame rate 30 fps Motion drive mechanism: A 6-axis robot arm with a plunger (corresponding to a pressing part) attached to the end of the arm. Motion conditions: Vertical pressing motion with a speed of 2 mm / s and a pressing depth of z = 9 mm. Measurement targets: Eight types of gel-like foods in cylindrical shapes, 20mm in diameter and 10mm in height.
[0055] The leave-one-out cross-validation method is used for estimation. In this method, given a total of M data points, a model is created using M-1 data points, and estimation is performed on the remaining 1 data point. This process is repeated M times.
[0056] Figure 7 shows the estimated results for the four textures. The horizontal axis represents the human sensory evaluation value n. i The vertical axis represents the estimated value. The estimated results for each gel-like food are plotted in eight different colors (A to H), with the average estimated value for each type plotted as a red square. The coefficient of determination, which represents the predictive performance of the model, is used as an indicator to compare the estimation accuracy. For all four textures, the coefficient of determination (R) 2 The large size of the ) suggests that it is possible to estimate texture with high accuracy.
[0057] (Embodiment 2) In Embodiment 2, the visual-tactile receiving portion 1 has a projection (P_i) formed on the second member 20. i is from 1 to n. One or more markers may be placed on the projection. Figure 3 shows an example of a surface projection and marker. The human tongue has several small projections called filiform papillae on the surface of its elastic tongue muscles. These filiform papillae are schematically reproduced by arranging projections on the surface. The projections on the surface reduce the pulling of adjacent markers, and because the markers can move more freely than in the planar arrangement of Embodiment 1, higher-resolution sensing is possible.
[0058] Other components, such as the motion driving means 30, evaluation unit 50, and output unit 60, are the same as in Embodiment 1.
[0059] (Example 6: Rigid Object) First component 10: Made of transparent, rigid acrylic, 4mm thick, 49mm x 49mm in size. Second component 20: Made of transparent, soft silicone rubber, with an elastic modulus of 55 kPa, a thickness of 15 mm, and dimensions of 49 mm x 49 mm. The perimeter is surrounded by a 3 mm thick PLA resin component. There are 225 protrusions in total, arranged in 15 rows and 15 columns, spaced 3.5 mm apart, and the shape of the protrusions is a hemisphere with a diameter of 3.5 mm. Markers (M_i): 225 black-painted glass spheres with a diameter of 1.0 mm, arranged in 15 rows and 15 columns, with a spacing of 3.5 mm between them. Image capture conditions: Resolution 1080 x 1080 pixels, frame rate 30 fps Motion drive mechanism: A 6-axis robotic arm with a plunger attached to the end of the arm. Motion conditions: Vertical translational motion after a depressure with a velocity of 2 mm / s and a depressure amount z = 9 mm. Measurement target: A cylindrical rigid object with a diameter of 20 mm and a height of 7 mm. The surface of the cylinder has irregularities at an angle θ, resulting in different surface textures. Figure 8A shows examples of four types of rigid objects with θ = {45, 90, 135, 180} degrees.
[0060] A rigid object is moved in a translational motion on the surface, and the marker displacement at this time is obtained. Figure 8B shows the marker displacement vectors and indentation amount z for θ = 180deg and 90deg for sensors without protrusions and sensors with protrusions. For a rigid object with θ = 180 degrees, the sensor without protrusions shows displacement changes according to the relative position between the central marker and the object. On the other hand, with the sensor with protrusions, the displacement is somewhat oscillatory, even though the surface of the rigid object is flat. This is presumed to be caused by the repeated pulling and release of the protrusions due to friction between the sensor's protrusions and the rigid object. For a rigid object with θ = 90 degrees, the marker displacement is more vibrational in the sensor without protrusions. In the sensor with protrusions, the markers within the contact surface between the sensor surface and the rigid object are displaced in various directions. This means that the protrusions reduce the tension on the soft second member 20, allowing each marker to move more freely.
[0061] Figure 8C shows the marker displacement, its average frequency f_ave, and average amplitude A_ave for four types of objects. For sensors without protrusions, differences in average frequency are observed depending on the surface texture of the object, but little difference is seen in average amplitude. On the other hand, sensors with protrusions can more clearly identify the surface texture of the object in both average frequency and average amplitude features. By reducing the friction between adjacent markers, the sensitivity to surface roughness of the object being measured was improved.
[0062] (Example 7: Flexible object) The measurement was performed on a cylindrical gel-like food product with a diameter of 20 mm and a height of 10 mm. The cylindrical surface had irregularities at an angle θ, and the measurement was carried out using four different shapes with the same θ = {45, 90, 135, 180} degrees as the rigid object. This is the same as in Example 6.
[0063] Figure 9A shows the marker displacement vectors and indentation amount z for sensors without protrusions and with protrusions at θ = 180deg and 90deg. It also shows the surface texture of the object. The trend was similar to that of the rigid object in Figure 8B. Figure 9B shows the marker displacement, its average frequency f_ave, and average amplitude A_ave for four types of gel-like food objects. With the sensor without protrusions, it was difficult to clearly distinguish between the types of objects by comparing either the average frequency or the average amplitude. On the other hand, with the sensor with protrusions, there was little difference in the average frequency, but a large difference in the average amplitude was observed, making identification possible. In the case of rigid objects, since the rigid object does not deform, the characteristics of the surface texture were reflected in the frequency of the marker displacement. On the other hand, in the case of flexible objects, the flexible object deforms due to the restoring force of the sensor, so the vibration of the marker is dominated by the vibration of the sensor itself, and no difference was observed in frequency, but a difference appeared in amplitude. Although the characteristic quantities that show differences change depending on the characteristics of the measured object, it was confirmed that the sensor with protrusions has improved discrimination performance for surface texture.
[0064] (Example 8: Identification and estimation of rigid surface) This was performed using both the sensor with protrusions and the sensor without protrusions as in Example 6. The markers were 1.0 mm in diameter black-painted glass spheres, with a total of 225 markers arranged in 15 rows and 15 columns, spaced 3.5 mm apart. The shape of the sensor protrusions was a hemisphere with a diameter of 3.5 mm. Figure 10A shows examples of rigid bodies with 42 patterns of convex surfaces. The height of the convex surface is set to 1.75 [mm], the same as the height of the apex of the sensor's protrusion. The diameter of the base circle is set to d=3.5 [mm] for the smallest diameter, and also to 2d and 3d for the other two diameters. Four textures are generated by rotating one object by π / 2 [rad], resulting in a total of 4 × 42 = 168 textures. The pressing and rotation of the object is performed as follows: First, the object is pressed down to the maximum pressing amount Z [mm]. After pressing is complete, the object is rotated by π / 2 [rad] at π / 8 [rad / s]. The moment the object touches the sensor is set to Z=0, and the maximum pressing amounts are set to two types: Z={0.5} and Z=1.0} [mm]. The pressing and rotation experiments under the above two conditions were performed five times for each texture. The shooting conditions are a resolution of 1920 x 1080 pixels and a frame rate of 30 fps.
[0065] The video footage acquired by the camera is analyzed to obtain the coordinates (x,y) of each marker based on its initial position. The coordinates of all markers from 50 frames, starting at random times during the rotation of the object being measured, are input to the ViViT learning model. The ViViT learning model consists of an encoder and a decoder. The encoder extracts a 256-dimensional feature vector from the input data, and the decoder, using transposed convolution, reconstructs it into a 6x6x3 size. The ViViT learning model outputs the probability that each of the 6x6 cells has a convex size of d, 2d, or 3d.
[0066] The most probable value is treated as the estimated value for each cell. The method for calculating the estimation accuracy α is described below. First, the correct labels are quantified as shown in Figure 10B. The texture is divided into 6x6 cells, and the protrusion size is stored in each cell. d is converted to 0, 2d to 1, 3d to 2, and so on. The estimation accuracy α for each texture is calculated by comparing the estimated value with the correct label for each cell. α = (number of correct cells for the estimated value) / (total number of cells = 36). This is done for all textures, and the average is taken as the sensor's estimation accuracy α. Two methods were used for training and validation. In validation method 1, of the five experiments conducted using 168 types of textures, four were used as training data and one as validation data. Since the training data and validation data include texture data under the same experimental conditions, the estimation accuracy is expected to be high. In Verification Method 2, the model was trained using 152 textures (38 types of measured objects), excluding the four rightmost textures (red frame) in the bottom row of Figure 10A. The 16 textures (4 types of measured objects) in the red frame were used as unknown texture patterns for verification. Unlike Verification Method 1, the texture patterns used in the verification data were not present in the training data, so the estimation accuracy was expected to be lower than in Verification Method 1. This analysis was performed for two different indentation conditions. The same experiment and analysis were also performed for a flat sensor (without protrusions), and the estimation accuracy of each sensor was compared. Training was performed with a learning rate of 0.3, a batch size of 16, and 2000 epochs.
[0067] Figure 10C(a) shows the sensor output image for the flat sensor at the top, and Figure 10C(b) shows the sensor output image for the protrusion sensor at the top. The results for indentation amounts Z=1.0[mm] and Z=0.5[mm] are shown. Regardless of the indentation amount, by looking at the marker displacement vector, it can be seen that the marker moves more with the protrusion sensor than with the conventional flat sensor. This indicates that the protrusion sensor responds more sensitively to the texture of the object. Also, when the indentation amount of the flat sensor is Z=0.5[mm], the marker displacement is small, and there is a high possibility that the texture of the object is not being captured. The texture estimation accuracy of the protrusion sensor and the flat sensor are compared. The estimation accuracy results for the flat sensor are shown at the bottom of Figure 10C(a) and the protrusion sensor at the bottom of Figure 10C(b). When the indentation amount of the flat sensor decreases from Z=1.0[mm] to Z=0.5[mm], the estimation accuracy drops significantly, from 17.7 points in verification method 1 to 22.9 points in verification method 2. In contrast, with the protrusion sensor, the difference in estimation accuracy remained within 5 points regardless of the amount of indentation, regardless of the verification method. This is presumed to be due to the fact that by providing a flexible group of protrusions containing a marker, the measurement part deformed without being pulled by fine textures due to the independence of the deformation of the protrusion group, and the sensitivity in the horizontal direction was increased.
[0068] (Example 9: Estimation of food texture (residue)) Measurement target: Nursing care food (commercially available products A, B, and C) Motion conditions: Vertical indentation with a velocity of 1 mm / s and an indentation amount z = 8 mm followed by rotation. After the indentation was completed, the object was rotated 2π [rad] at a speed of π / 8 [rad / s], and then rotated -2π [rad] for 5 sets. Figure 11 shows the state of the marker displacement vector obtained by analyzing the image data acquired from the initial state up to five reciprocating rotations.
[0069] Eight subjects consumed the food and evaluated the residue left behind when crushed with the tongue on a 7-point scale (1: no residue at all, 4: neither, 7: very residue), and the average value was calculated. As shown in Figure 11, samples with higher "residue" scores showed a greater number of markers with large displacement vectors because unbreakable particles remained on the sensor even after repeated pressing and rotating. The magnitude and number of marker displacement vectors when the object was placed on the sensor and repeatedly compressed and rotated corresponded to the "residue" when crushed with the tongue. This is shown in Table 1.
[0070] [Table 1]
[0071] (Example 10: Estimation of stickiness) (1) Sensor without protrusions First component 10: Transparent, rigid acrylic, 4mm thick, 40mm x 40mm in size. Second component 20: Made of transparent, soft silicone rubber, with an elastic modulus of 98 kPa, a thickness of 10 mm, and dimensions of 40 mm x 40 mm, surrounded by a 4 mm thick PLA resin component. Markers (M_i): 442 black-painted glass spheres, 0.6 mm in diameter, arranged in 21 rows and 21 columns, spaced 2 mm apart. (2) Sensor with protrusions First component 10: Made of transparent, rigid acrylic, 4mm thick, 49mm x 49mm in size. Second component 20: Made of transparent, soft silicone rubber, with an elastic modulus of 55 kPa, a thickness of 15 mm, and dimensions of 49 mm x 49 mm. The perimeter is surrounded by a 3 mm thick PLA resin component. There are a total of 225 protrusions arranged in 15 rows and 15 columns, spaced 3.5 mm apart, and the shape of the protrusions is a hemisphere with a diameter of 3.5 mm. Markers (M_i): 225 black-painted glass spheres with a diameter of 1.0 mm, arranged in 15 rows and 15 columns, with a spacing of 3.5 mm between them. Image capture conditions: Resolution 1080 x 1080 pixels, frame rate 30 fps Motion drive mechanism: A 6-axis robotic arm with a plunger attached to the end of the arm. Motion conditions: Vertical indentation with a speed of 2 mm / s and indentation amount z = 8 mm followed by rotational motion. After the pressing was complete, the object was rotated 2π rad at π / 8 rad / s, and then rotated -2π rad five times. Measurement subjects: Three types of carrageenan-containing gel-like foods (#7-1, C1, C3) in cylindrical shapes with a diameter of 20 mm and a height of 10 mm. The degree of stickiness was in the order of C1 > C3 > #7-1. Measurements were taken for 8 times for C1 and C3, and 4 times for #7-1.
[0072] The results of the sensory evaluation test for stickiness are shown below. Five subjects consumed the food and evaluated its stickiness when crushed with their tongues on a 7-point scale (1: not sticky at all, 4: neither sticky nor sticky, 7: very sticky). The average value was calculated. The results are shown in Table 2.
[0073] [Table 2]
[0074] Figure 12 shows the relationship of the time course of the sum of marker displacements [px]. Figure 12(a) shows the results for a sensor without protrusions, and (b) shows the results for a sensor with protrusions. The average values of the three types are shown normalized by the maximum value. The sum of marker displacements showed a repeating convex waveform with short intervals between valleys after the peak. It can be seen that the sensor with protrusions has higher discriminative ability than the sensor without protrusions. Furthermore, in C1, which has a stronger sticky feeling, this repeating waveform is positioned higher than the others, meaning that the sum of marker displacements was larger. From the above, it was found that the sticky feeling can be estimated from the time course of the sum of marker displacements and the shape of the waveform. [Explanation of Symbols]
[0075] A1 Visual and tactile sensor device 1. Visual and tactile receptors 10 First component 20 Second component 30. Motion driving means 50 Evaluation Department 51 Marker displacement acquisition unit 52 Learning Models 60 Output section M Marker P protrusion
Claims
1. A tactile and visual sensory receptor portion that is partially or entirely light-transmitting in the thickness direction, A plurality of markers are arranged at predetermined positions inside one surface in the thickness direction of the visual-tactile receiving portion, A motion driving means that brings the object to be measured into contact with the tactile receiving portion on the marker side and performs a predetermined action, While the motion driving means is causing the object to be measured to perform the predetermined action, an imaging unit captures the positional state of the marker from the optic-tactile receiving part side that is not in contact with the object to be measured, The system includes an evaluation unit that analyzes image data captured by the imaging unit and evaluates one or more of the shape, size, distribution, hardness, and texture of the object to be measured. Visual and tactile sensor device.
2. The visual-tactile sensor device according to claim 1, wherein the visual-tactile receiving portion has a flat surface on one side.
3. The optic and tactile receiving portion has a plurality of protrusions on one surface side, and one or more markers are provided on each of the protrusions. The visual and tactile sensor device according to claim 1.
4. The evaluation unit described above, A marker displacement acquisition unit obtains marker displacement information for each of the markers from the acquired time-series image data, The system includes a first estimation unit that evaluates one or more of the shape, size, distribution, hardness, and texture of the object to be measured based on the marker displacement information obtained by the marker displacement acquisition unit. The visual and tactile sensor device according to claim 1.
5. The evaluation unit described above, The system further includes a feature calculation unit that calculates the feature quantities of the marker displacement obtained by the marker displacement acquisition unit, The first estimation unit evaluates one or more of the shape, size, distribution, hardness, and texture of the object to be measured based on the marker displacement information and / or the characteristic quantities of the marker displacement. The visual and tactile sensor device according to claim 4.
6. The evaluation unit described above, A storage device that stores a learning model generated by intelligent information processing technology using one or more data selected from marker displacement information and marker displacement feature quantities obtained from image data captured by the imaging unit of a visual-tactile sensor device, using a measurement target for which one or more of the following are known: shape, size, distribution, hardness, and texture, and one or more of the data selected from shape (surface characteristics), size, distribution, hardness, and texture as training data, The system includes a second estimation unit that uses the learning model to identify and classify one or more of the shape, size, distribution, hardness, and texture, using marker displacement information and one or more data selected from the marker displacement features obtained from the image data captured by the imaging unit as input data. The visual and tactile sensor device according to claim 1.
7. A method for evaluating texture using the visual and tactile sensor device described in claim 1, From the acquired time-series image data, marker displacement information and / or marker displacement feature quantities are obtained for each marker, and based on the marker displacement information and / or marker displacement feature quantities, one or more of the shape (surface properties), size, distribution, hardness, and texture of the object to be measured are evaluated in the first evaluation step, and / or, Using a measurement target for which one or more of the following characteristics are known, the method includes a second evaluation step in which, using one or more data selected from marker displacement information and marker displacement feature quantities obtained from image data captured by the imaging unit of a visual-tactile sensor device, and one or more data selected from the shape, size, distribution, hardness, and texture as training data, one or more data selected from the marker displacement information and marker displacement feature quantities obtained from the captured image data is input to a learning model generated by intelligent information processing technology, and one or more data selected from the marker displacement information and marker displacement feature quantities obtained from the captured image data is input to identify and classify one or more of the shape, size, distribution, hardness, and texture. Method for evaluating texture.
8. It is a texture evaluation program, A program that enables the texture evaluation method described in claim 7 using at least one processor.
9. At least one processor, The processor includes a memory for storing instructions that can be executed by the processor, The processor is an information processing device that realizes the steps of the texture evaluation method described in claim 7 by executing an executable instruction.