Standardized tongue picture acquisition device and detection method

By designing a scalable tongue image acquisition device and image processing algorithm, the problems of large device size and the influence of environmental factors were solved, and the rapid and accurate acquisition and analysis of tongue image characteristics were achieved.

CN120661082APending Publication Date: 2025-09-19GUANGZHOU UNIVERSITY OF CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510671882.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing tongue image collection devices are large in size, inconvenient to use, require high professionalism, environmental factors affect the consistency and accuracy of feature recognition, and background information interference is serious.

Method used

A tongue image acquisition device with a retractable L-shaped shell structure was designed. The image processing algorithm of Fourier transform and Haar wavelet transform was combined to remove background information and realize rapid collection and analysis of tongue images.

Benefits of technology

The device is small and portable and can adapt to different environments. The image processing algorithm can quickly remove the background and improve the accuracy and consistency of tongue feature recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120661082A_ABST
    Figure CN120661082A_ABST
Patent Text Reader

Abstract

The invention provides a standardized tongue picture collecting device which comprises a shell, when observed from the side face, the shell is in an L shape and is divided into a horizontal part and a vertical part, the horizontal part comprises a first shell and a second shell arranged at one end of the first shell, and the first shell is telescopically connected relative to the second shell; the vertical part comprises a third shell and a fourth shell connected with the third shell, and the third shell is telescopically connected with the fourth shell; and the second shell is connected with the third shell.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of tongue image detection, and in particular to a standardized tongue image collection device and detection method. Background Art

[0002] The tongue image acquisition device is a medical auxiliary equipment that uses cameras and computer technology to collect and process tongue image feature information. By identifying details such as the tongue's color, shape, and tongue coating, it uploads the information to software for tongue image feature analysis to achieve objective quantification of traditional Chinese medicine tongue diagnosis. It can be used in auxiliary diagnosis scenarios at home or in medical institutions.

[0003] The more common tongue image collection devices on the market currently have the following three problems:

[0004] 1. The machine is large in size, inconvenient to use, and requires high professionalism of the user

[0005] 2. Due to environmental factors such as light, distance, and collection angle, tongue feature recognition lacks consistency and accuracy.

[0006] 3. In the collected images, in addition to the key tongue part, there are often background areas such as lips, facial skin, and teeth. However, this background information does not have significant reference value in tongue feature analysis, causing unnecessary interference. Summary of the Invention

[0007] The main purpose of the present invention is to provide a standardized tongue image collection device to solve the above technical problems.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is: a standardized tongue image collection device includes a shell. When viewed from the side, the shell is L-shaped, consisting of a horizontal part and a vertical part. The horizontal part includes a first shell and a second shell arranged at one end of the first shell, and the first shell is telescopically connected to the second shell; the vertical part includes a third shell and a fourth shell connected to the third shell, and the third shell and the fourth shell are telescopically connected; the second shell is connected to the third shell.

[0009] Preferably, the first shell and the second shell are connected by a telescopic mechanism, and the third shell and the fourth shell are also connected by the same telescopic mechanism; the telescopic mechanism includes a connecting block, and a plurality of receiving slots are provided on one side of the connecting block, and a button is provided on at least one of the first shell and the second shell and on at least one of the third shell and the fourth shell, and the button can cooperate with one of the receiving slots; a spring is provided between the button and the corresponding shell, and the spring is always in a compressed state, and the spring can push the button into the receiving slot; a strip groove is provided on the shell and at a position corresponding to the button, and the button extends through the strip groove to the outside of the shell, and the strip groove extends in the movable direction of the button, and the button can be disengaged from the corresponding receiving slot by manually operating the button, thereby realizing the telescopic extension of the corresponding shell.

[0010] Preferably, the connecting block is a hollow structure, which is conductive along the direction of extension and contraction. A sliding rail is provided on the inner wall of one side thereof, and two groups of sliding plates are provided on the sliding rail to match it. One group of sliding plates is connected to the outer shell on one side, and the other group of sliding plates is connected to the outer shell on the other side.

[0011] Preferably, it also includes a forehead locator, one end of which extends from the end of the first shell and can start to move relative to the first shell. The inner end of the forehead locator is provided with a button mechanism, and the button mechanism includes a button cover, a push rod, a spring and a push core. The button cover is provided at the inner end of the forehead locator, and the end of the button cover away from the forehead locator is provided with a first tooth, and the end of the push rod facing the button cover is provided with a second tooth. The first tooth and the second tooth are engaged with each other. By pressing the forehead locator, the push rod can be moved in the axial direction to move the push core between the first position and the second position.

[0012] Preferably, the pushing core is arranged at the end of the push rod away from the key cover, the spring is sleeved on the pushing core, one end of the spring abuts against the push rod, and the other end abuts against the inner wall of the first shell, and is always in a compressed state.

[0013] Preferably, a baffle is provided in the first shell and on the side of the pushing core facing away from the push rod. When the pushing core is in the first position, the pushing core rests against the baffle, at which time the current is conducted and the camera enters the standby state; when the pushing core is in the second position, the pushing core is disengaged from the baffle and the circuit is disconnected.

[0014] Preferably, a push rod, an anode wire, a wireless transmitter, a circuit board and a main control module arranged on the circuit board are arranged in the second shell, and a display screen and a battery box are arranged on the shell of the second shell. The battery box is used to place batteries, and a shooting button is arranged on the lower side of the battery box.

[0015] Preferably, a camera is provided on the front side of the fourth housing for photographing the tongue, and relevant data of the tongue image can be obtained by processing the photographed image.

[0016] The present invention also provides a method for processing images captured by the above-mentioned device, which specifically includes the following steps:

[0017] Step 1: Perform Fourier transform on the source image and the target image to obtain the frequency distribution of the source image and the target image;

[0018] Step 2: replace the low-frequency component of the source image with the low-frequency component of the target image amplitude without changing the phase component of the source image;

[0019] Step 3: Use inverse Fourier transform to convert the frequency domain obtained in step 2 into the spatial domain to obtain a spatial domain image;

[0020] Step 4: Process the spatial domain image obtained in step 3 using a downsampling module based on Haar wavelet transform. The downsampling module based on Haar wavelet transform includes a lossless feature encoding module and a feature representation learning module. The lossless feature encoding module uses Haar wavelet transform to reduce the spatial resolution of the feature space graph while maintaining all the information of the spatial domain image.

[0021] Step 5: Use dilated spatial pyramid pooling to perform multi-scale feature extraction on the image output in step 4 to obtain a feature map;

[0022] Step 6: Use multiple upsampling layers CCU for upsampling. The input of the first upsampling layer is the feature map obtained by dilated spatial pyramid pooling of the output of the last downsampling layer. The input of the upsampling layers of other layers is the concatenation of the feature maps obtained by dilated spatial pyramid pooling of the output of the previous upsampling layer and the output of the corresponding downsampling layer. The final output of the upsampling layer is a two-class mask.

[0023] Step 7: Superimpose the binary classification mask with the source image to obtain the tongue image with the background removed, which is the tongue image.

[0024] Preferably, the Fourier transform formula in step 1 is as follows:

[0025]

[0026] M×N represents the size of the source image or the target image, f(x,y) is the pixel of the source image or the target image at point (x,y), i 2 =-1.

[0027] Preferably, the Fourier inverse transform formula in step 3 is as follows:

[0028]

[0029] F -1 is the inverse Fourier transform that maps the frequency domain back to the spatial domain, F A and F p are the amplitude and phase components of the Fourier transform of the tongue image, is the low-frequency part of the target image amplitude, is the high-frequency part of the target image amplitude.

[0030] Preferably, step 4 specifically includes the following steps:

[0031] Step 41: Input the spatial domain image into a lossless feature coding module, where a wavelet transform layer (HWT) in the lossless feature coding module reduces the spatial resolution of the spatial domain image and retains all feature information of the spatial domain image;

[0032] The wavelet basis function and scaling function of the one-dimensional Haar transform of level 1 can be defined as follows:

[0033]

[0034] Among them, φ i,j (x) is defined as:

[0035]

[0036] In this case, the parameters j and k are expressed as the order of the Haar basis function, and φ 0,0 (x) is defined as:

[0037]

[0038] Therefore, the 1st-level Haar transform can be expressed using the 0th-level Haar basis function:

[0039]

[0040] Step 42: The feature representation learning module performs redundancy filtering on the image output by the lossless feature encoding module.

[0041] Preferably, the step 5 specifically includes the following steps:

[0042] Step 51: Fold the feature map: Assume that the size of the feature map is H×W×C, and use a×a window with a sliding step size of a on the feature map, where a is smaller than H, W, and C. Collect a×a×C feature points in each window and stack them according to the channel direction to obtain a H / a×W / a×a 2 Feature map;

[0043] Step 52: Perform convolution operations using dilated convolutions with different expansion rates and sizes to obtain corresponding feature maps.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] The device of the present invention is small in size and easy to carry; the shell is retractable and can be adjusted according to different situations; and an image processing algorithm is integrated inside, and after the shooting is completed, the tongue image can be quickly obtained through image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a perspective view of the present invention;

[0047] Figure 2 It is an internal diagram of the present invention;

[0048] Figure 3 is another internal view of the present invention;

[0049] Figure 4 is a schematic diagram of the image processing method of the present invention. DETAILED DESCRIPTION

[0050] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.

[0051] Example 1

[0052] A standardized tongue image collection device includes a housing that, when viewed from the side, is L-shaped. For ease of description, the housing is referred to as the horizontal portion and the vertical portion, respectively. The vertical portion also serves as a handle for gripping. The horizontal portion comprises a first housing 2 and a second housing 5 disposed at one end of the first housing 2. The first housing 2 is telescopically connected to the second housing 5 to adjust the length of the horizontal portion. The vertical portion comprises a third housing 6 and a fourth housing 8 connected to the third housing 6. The third and fourth housings 6 and 8 are telescopically connected to adjust the length of the vertical portion. The second housing 5 is connected to the third housing 6. Specifically, a telescopic mechanism 18 connects the first and second housings 2 and 5, and the same telescopic mechanism 18 connects the third and fourth housings 6 and 8. The telescopic mechanism 18 comprises a connecting block with multiple receiving slots disposed on one side of the connecting block. A button 3 is disposed on at least one of the first and second housings 2 and 5, and on at least one of the third and fourth housings 6 and 8. The button 3 is capable of engaging with one of the receiving slots. A spring is installed between button 3 and the corresponding housing. The spring is always compressed, pushing the button into the receiving slot. A strip groove is provided on the housing at a position corresponding to the button. The button 3 extends through the strip groove to the outside of the housing. The strip groove extends in the direction of button 3's movement. Manually operating the button 3 can remove the button 3 from the corresponding receiving slot, thereby achieving the corresponding housing extension and retraction.

[0053] Furthermore, the connecting block is a hollow structure, which is conductive along the direction of extension and contraction. A slide rail 17 is provided on the inner wall of one side thereof, and two groups of sliding plates 16 are provided on the slide rail 17 to cooperate with it. One group of sliding plates 16 is connected to the outer shell on one side, and the other group of sliding plates 16 is connected to the outer shell on the other side. The slide rail 17 and the sliding plates 16 can enhance the stability of the extension and contraction of the first outer shell 2 and the second outer shell 5.

[0054] The device also includes a forehead locator 1, one end of which extends from the end of the first housing 2 and can begin to move relative to the first housing 2. A key mechanism 28 is provided at the inner end of the forehead locator 1. The key mechanism 28 includes a key cover 11, a push rod 12, a spring 13, and a push core 14. The key cover 11 is provided at the inner end of the forehead locator 1. The end of the key cover 11 away from the forehead locator 1 is provided with a first tooth. The end of the push rod 12 facing the key cover 11 is provided with a second tooth. The first and second teeth engage with each other. The principle of cooperation between the push rod 12 and the key cover 11 is the same as that of an automatic ballpoint pen. By pressing the forehead locator 1, the push rod 12 can be moved axially, so that the push core 14 moves between the first position and the second position.

[0055] The pushing core 14 is arranged at the end of the push rod 12 facing away from the button cover 11, and the spring 13 is sleeved on the pushing core 14. One end of the spring 13 rests on the push rod 12, and the other end rests on the inner wall of the first shell 2, and is always in a compressed state, so that the push rod 12 and the button cover 11 always rest against each other and ensure that the forehead locator 1 does not move by itself.

[0056] A baffle 15 is provided inside the first shell 2 and on the side of the pushing core 14 facing away from the push rod 12. When the pushing core 14 is in the first position, the pushing core 14 rests on the baffle 15, at which time the current is conducted and the camera enters the standby state; when the pushing core 14 is in the second position, the pushing core 14 is disengaged from the baffle 15 and the circuit is disconnected.

[0057] A wireless transmitter 29, a circuit board 22, and a main control module 23 arranged on the circuit board 22 are arranged in the second shell 5. A display screen 20 and a battery box 24 are arranged on the shell of the second shell 5. The battery box 24 is used to place batteries. A shooting button 9 is arranged on the lower side of the battery box 24.

[0058] A camera 7 is provided on the front side of the fourth housing 8 for photographing the tongue. The photographed image can be processed to obtain relevant data of the tongue image.

[0059] Example 2

[0060] This embodiment is a method for processing the picture taken in the first embodiment, which specifically includes the following steps:

[0061] Step 1: Perform Fourier transform on the source image and the target image to obtain the frequency distribution of the source image and the target image. The Fourier transform formula is as follows:

[0062]

[0063] M×N represents the size of the source image or the target image, f(x,y) is the pixel of the source image or the target image at point (x,y), i 2 =-1;

[0064] Step 2: replace the low-frequency component of the source image with the low-frequency component of the target image amplitude without changing the phase component of the source image;

[0065] Step 3: Use the inverse Fourier transform to convert the frequency domain obtained in step 2 to the spatial domain to obtain the spatial domain image. The inverse Fourier transform formula is as follows:

[0066]

[0067] F -1 is the inverse Fourier transform that maps the frequency domain back to the spatial domain, FA and F p are the amplitude and phase components of the Fourier transform of the tongue image, is the low-frequency part of the target image amplitude, It is the high frequency part of the target image amplitude.

[0068] Step 4: Use a downsampling module (HWD) based on Haar wavelet transform to process the spatial domain image obtained in step 13. The downsampling module based on Haar wavelet transform includes a lossless feature encoding module and a feature representation learning module. The lossless feature encoding module uses Haar wavelet transform to reduce the spatial resolution of the feature space map while maintaining all information; the feature representation learning module consists of a standard 1×1 convolution layer, a batch normalization layer and a ReLU activation function to extract discriminative features. The feature representation learning module is used to adjust the number of channels of the feature map to adapt to subsequent layers and filter redundant information. The processing specifically includes the following steps:

[0069] Step 41: Input the spatial domain image into a lossless feature coding module, where a wavelet transform layer (HWT) in the lossless feature coding module reduces the spatial resolution of the spatial domain image and retains all feature information of the spatial domain image;

[0070] The wavelet basis function and scaling function of the one-dimensional Haar transform of level 1 can be defined as follows:

[0071]

[0072] Among them, φ i,j (x) is defined as:

[0073]

[0074] In this case, the parameters j and k are expressed as the scale factor and translation factor of the Haar basis function, respectively, where the scale factor corresponds to frequency and the translation factor corresponds to time, and φ 0,0 (x) is defined as:

[0075]

[0076] Therefore, the 1st-level Haar transform can be expressed using the 0th-level Haar basis function:

[0077]

[0078] This means that a signal of length L can be split into two parts of length L / 2, which can be interpreted as low-pass and high-pass decomposition filters, respectively. When the Haar wavelet transform is applied to a two-dimensional signal (such as a grayscale image), four components are produced, each with half the spatial resolution of the original signal.

[0079] Step 42: The feature representation learning module performs redundancy filtering on the image output by the lossless feature encoding module to obtain a downsampled image.

[0080] Step 43, repeatedly performing steps 41 and 42 for a predetermined number of times, performing multiple downsampling on the spatial domain image, and obtaining multiple downsampled images;

[0081] Step 5: Use Atrous Spatial Pyramid Pooling (ASPP) to perform multi-scale feature extraction on the downsampled image output in step 4 to obtain a feature map, which specifically includes the following steps:

[0082] Step 51: Fold the feature map: Assume that the size of the feature map is H×W×C, and use a sliding port on the feature map. a is smaller than H, W, and C. In each window, collect feature points of a×a×C and stack them according to the channel direction to obtain a H / a×W / a×a 2 Feature map;

[0083] Step 52: Perform convolution operation using dilated convolutions with different expansion rates and sizes to obtain corresponding feature maps.

[0084] Step 6: Use multiple upsampling layers CCU for upsampling. The input of the first upsampling layer is the feature map obtained by dilated spatial pyramid pooling of the output of the last downsampling layer. The input of the upsampling layers of other layers is the concatenation of the feature maps obtained by dilated spatial pyramid pooling of the output of the previous upsampling layer and the output of the corresponding downsampling layer. The final output of the upsampling layer is a two-class mask.

[0085] Step 7: Superimpose the binary classification mask with the source image to obtain the tongue image with the background removed, which is the tongue image.

[0086] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A standardized tongue image collection device, characterized in that: The device comprises a housing, which is L-shaped when viewed from the side and comprises a horizontal portion and a vertical portion, wherein the horizontal portion comprises a first housing and a second housing provided at one end of the first housing, the first housing being telescopically connected to the second housing; and the vertical portion comprises a third housing and a fourth housing connected to the third housing, the third housing and the fourth housing being telescopically connected. The second housing is connected to the third housing.

2. The standardized tongue image collection device according to claim 1, characterized in that: The first shell and the second shell are connected by a telescopic mechanism, and the third shell and the fourth shell are also connected by the same telescopic mechanism; the telescopic mechanism includes a connecting block, and a plurality of receiving slots are provided on one side of the connecting block. A button is provided on at least one of the first shell and the second shell and on at least one of the third shell and the fourth shell, and the button can cooperate with one of the receiving slots; a spring is provided between the button and the corresponding shell, and the spring is always in a compressed state, and the spring can push the button into the receiving slot; a strip groove is provided on the shell and at a position corresponding to the button, and the button extends through the strip groove to the outside of the shell, and the strip groove extends in the movable direction of the button. The button can be disengaged from the corresponding receiving slot by manually operating the button, thereby realizing the telescopic movement of the corresponding shell.

3. The standardized tongue image collection device according to claim 2, characterized in that: The connecting block is a hollow structure and is conductive along the direction of extension and contraction. A sliding rail is provided on the inner wall of one side of the connecting block. Two sets of sliding plates are provided on the sliding rail to match the connecting block. One set of sliding plates is connected to the outer shell on one side, and the other set of sliding plates is connected to the outer shell on the other side.

4. The standardized tongue image collection device according to claim 3, characterized in that: It also includes a forehead locator, one end of which extends from the end of the first shell and can start to move relative to the first shell. A button mechanism is provided at the inner end of the forehead locator, and the button mechanism includes a button cover, a push rod, a spring and a push core. The button cover is provided at the inner end of the forehead locator, and the end of the button cover away from the forehead locator is provided with a first tooth, and the end of the push rod facing the button cover is provided with a second tooth. The first tooth and the second tooth are engaged with each other. By pressing the forehead locator, the push rod can be moved in the axial direction to move the push core between the first position and the second position.

5. The standardized tongue image collection device according to claim 4, characterized in that: The pushing core is arranged at the end of the push rod away from the key cover, the spring is sleeved on the pushing core, one end of the spring abuts against the push rod, the other end abuts against the inner wall of the first shell, and is always in a compressed state.

6. The standardized tongue image collection device according to claim 5, characterized in that: A baffle is provided inside the first shell and on the side of the pushing core facing away from the push rod. When the pushing core is in the first position, the pushing core rests against the baffle, at which time the current is conducted and the camera enters the standby state; when the pushing core is in the second position, the pushing core is disengaged from the baffle and the circuit is disconnected.

7. The standardized tongue image collection device according to claim 6, characterized in that: A camera is provided on the front side of the fourth housing for photographing the tongue. The photographed image can be processed to obtain relevant data of the tongue image.

8. A method for processing an image captured by the standardized tongue image acquisition device according to any one of claims 1 to 7, comprising the following steps: Step 1: Perform Fourier transform on the source image and the target image to obtain the frequency distribution of the source image and the target image; Step 2: replace the low-frequency component of the source image with the low-frequency component of the target image amplitude without changing the phase component of the source image; Step 3: Use inverse Fourier transform to convert the frequency domain obtained in step 2 into the spatial domain to obtain a spatial domain image; Step 4: Process the spatial domain image obtained in step 3 using a downsampling module based on Haar wavelet transform. The downsampling module based on Haar wavelet transform includes a lossless feature encoding module and a feature representation learning module. The lossless feature encoding module uses Haar wavelet transform to reduce the spatial resolution of the feature space graph while maintaining all the information of the spatial domain image. Step 5: Use dilated spatial pyramid pooling to perform multi-scale feature extraction on the image output in step 4 to obtain a feature map; Step 6: Use multiple upsampling layers CCU for upsampling. The input of the first upsampling layer is the feature map obtained by dilated spatial pyramid pooling of the output of the last downsampling layer. The input of the upsampling layers of other layers is the concatenation of the feature maps obtained by dilated spatial pyramid pooling of the output of the previous upsampling layer and the output of the corresponding downsampling layer. The final output of the upsampling layer is a two-class mask. Step 7: Superimpose the binary classification mask with the source image to obtain the tongue image with the background removed, which is the tongue image.

9. The processing method according to claim 8, characterized in that: The Fourier transform formula in step 1 is as follows: M×N represents the size of the source image or the target image, f(x,y) is the pixel of the source image or the target image at point (x,y), i 2 =-1.

10. The processing method according to claim 8, characterized in that: The step 4 specifically includes the following steps: Step 41: input the spatial domain image into a lossless feature coding module, wherein a wavelet transform layer in the lossless feature coding module reduces the spatial resolution of the spatial domain image and retains all feature information of the spatial domain image; The wavelet basis function and scaling function of the one-dimensional Haar transform of level 1 can be defined as follows: Among them, φ i,j (x) is defined as: Parameters j and k represent the scale factor and translation factor of the Haar basis function, respectively, and φ 0,0 (x) is defined as: Therefore, the 1st-level Haar transform can be represented by the 0th-level Haar basis function: