Tongue picture detection system and method based on standard color reference
By combining an LED array module, a diffuse polarization module, a light source stabilization module, and a color reference module, the problem of acquiring high-quality images in different environments for tongue image detection equipment is solved, achieving a balance between portability and color correction, and improving the accuracy of detection results.
Patent Information
- Application Number
- CN202510945998.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-24
AI Technical Summary
Existing tongue image detection equipment cannot acquire high-quality tongue images in various environments, resulting in inaccurate detection results. Furthermore, portability and color correction cannot be simultaneously achieved.
It employs an LED array module, a diffuse polarization module, a light source stabilization module, and a color reference module. Through cross illumination, diffuse and polarization technologies, combined with constant current drive and color difference calibration, it achieves stable light source and automatic color correction.
It improves the clarity and accuracy of tongue images, ensuring high-quality image acquisition in different environments, and achieves a balance between portability and color correction.
Smart Images

Figure CN120827342A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a tongue image detection system and method based on standard color reference. BACKGROUND
[0002] With the development of digital image analysis technology, the objectification and intelligentization of tongue image detection have become an important research direction.
[0003] However, there is currently a lack of a device that is extremely portable, always available and automatically deployed, integrates standard color reference and automatically color corrects in the field of tongue image detection. Existing portable devices usually sacrifice color quality, while systems that focus on color correction are often not portable enough.
[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide a tongue image detection system and method based on standard color reference, aiming to solve the problems of color distortion and inconsistent quality of tongue image in the prior art, which cannot collect high-quality tongue images in various environments, resulting in inaccurate tongue image detection results.
[0006] To achieve the above purpose, the present application provides a tongue image detection system based on standard color reference, wherein the tongue image detection system based on standard color reference comprises:
[0007] The LED array module is used to control the cross-illumination of a plurality of LED lights and convert the illumination light of each LED light into corresponding unidirectional soft light through a ring-shaped diffusion cover;
[0008] The diffuse polarization module is used to convert all the unidirectional soft light into combined soft light, filter the combined soft light through a plurality of linear polarizing plates, and obtain diffuse light;
[0009] The light source stabilization module is used to control the actual current passing through all the LED lights in real time and control the camera to delay exposure for shooting under the irradiation of the diffuse light to obtain a tongue image;
[0010] The color reference module is used to calibrate the tongue image multiple times for color difference and judge the calibrated image to output a detection result.
[0011] The tongue image detection system based on standard color reference, wherein the LED array module comprises an LED light sub-module and a ring-shaped diffusion cover;
[0012] The LED lamp sub-module comprises a plurality of LED lamps with preset color rendering indexes, all of which are arranged in a ring shape on a ring-shaped printed circuit board at a preset inclination angle, so that the optical axis of each LED lamp points to the front of the central axis of the ring-shaped printed circuit board.
[0013] The ring-shaped diffusion cover covers the front of the LED lamp sub-module at a preset distance from all the LED lamps, so as to soften the illumination light of all the LED lamps and obtain corresponding unidirectional soft light.
[0014] The tongue image detection system based on standard color reference, wherein the diffusion polarization module comprises a diffusion sub-module, a first linear polarizer and a second linear polarizer.
[0015] The diffusion sub-module is used for scattering the unidirectional soft light to obtain corresponding area light sources, and a combined soft light is obtained according to all the area light sources.
[0016] The first linear polarizer is arranged in front of the ring-shaped diffusion cover and is used for preliminarily filtering the combined soft light to obtain vertical light in the combined soft light.
[0017] The second linear polarizer is perpendicular to the first linear polarizer and is arranged in front of the camera and is used for filtering reflected light to obtain diffusion light.
[0018] The reflected light represents the light reflected after the vertical light irradiates the target tongue image.
[0019] The tongue image detection system based on standard color reference, wherein the light source stabilizing module comprises a current control sub-module and a shooting sub-module.
[0020] The current control sub-module is used for controlling the actual current flowing through all the LED lamps in real time through a feedback control loop, so that the actual current is kept at a preset current value.
[0021] The shooting sub-module is used for lighting all the LED lamps at the actual current for a preset time through a constant current driving chip, and controls the camera to shoot when the PN junction of all the LED lamps reaches an equilibrium point to obtain a tongue image.
[0022] The tongue image detection system based on standard color reference, wherein the color reference module comprises a global correction sub-module, a nonlinear calibration sub-module and a color judgment sub-module.
[0023] The global correction sub-module is used for performing white balance processing on the tongue image through a first color block set in advance to obtain a global calibration graph.
[0024] The nonlinear calibration submodule is configured to perform color calibration on multiple regions of the global calibration image by using a preset second color block, a third color block and a fourth color block, respectively, to obtain a target tongue image;
[0025] The color judgment submodule is configured to detect the target tongue image according to the first color block, the third color block and the fourth color block, respectively, to obtain a color difference detection result of the target tongue image.
[0026] The tongue image detection system based on the standard color reference, wherein the global correction submodule comprises a positioning unit, a coefficient calculation unit and a correction unit;
[0027] The positioning unit is configured to determine an accurate region of a neutral gray block in the tongue image according to the first color block;
[0028] The coefficient calculation unit is configured to calculate a first color channel gain, a second color channel gain and a third color channel gain in the accurate region;
[0029] The correction unit is configured to perform traversal update on each pixel in the tongue image according to the first color channel gain, the second color channel gain and the third color channel gain, to generate a corrected global calibration image.
[0030] The tongue image detection system based on the standard color reference, wherein the nonlinear calibration submodule comprises a color card positioning unit and a color calibration unit;
[0031] The color card positioning unit is configured to mark each region to obtain a corresponding rectangular contour, perform perspective transformation on all the rectangular contours to obtain a corresponding square contour, read a binary matrix code in each square contour to obtain a corresponding binary number, and query the binary number in a preset dictionary to obtain an ID number and multiple pixel coordinate points corresponding to each square contour;
[0032] The color calibration unit is configured to define multiple target coordinate points, input all the pixel coordinate points and all the target coordinate points into a matrix transformation model, output a transformation matrix, inversely calculate all the pixel points in the global calibration image by using the transformation matrix, obtain a color of each pixel point in the global calibration image, and fill the color of each pixel point into a new image to obtain a target tongue image.
[0033] In addition, the tongue image detection system based on the standard color reference also provides a tongue image detection method based on the standard color reference, wherein the tongue image detection method based on the standard color reference comprises:
[0034] The LED array module controls multiple LED lamps to cross-illuminate, and converts the illumination light of each LED lamp into corresponding unidirectional soft light through a ring-shaped diffusion cover;
[0035] The diffusion polarization module converts all the unidirectional soft light into combined soft light, filters the combined soft light through multiple linear polarizers, and obtains diffusion light;
[0036] The light source stabilization module controls the actual current passing through all the LED lamps in real time, controls the camera to delay exposure, and performs shooting under the illumination of the diffusion light, and obtains a tongue image;
[0037] The color reference module performs multiple chromatic aberration calibrations on the tongue image, judges the calibrated image, and outputs a detection result.
[0038] The tongue image detection method based on the standard color reference of the tongue image detection system based on the standard color reference, wherein the judging the calibrated image and outputting the detection result specifically includes:
[0039] The chromatic aberration judging unit judges the gradient positions of the colors of all regions of the target tongue image in the first color block, the third color block and the fourth color block, and when the chromatic aberration detection result meets the preset standard, performs segmentation processing on the target tongue image to obtain a tongue body region;
[0040] The feature extraction unit extracts features of the target tongue image according to the tongue body region, and obtains color features, texture features, morphological features and sublingual collateral features of the target tongue image;
[0041] The detection unit inputs the color features, the texture features, the morphological features and the sublingual collateral features into a constructed multi-modal model for detection, and outputs an abnormal detection result of the target tongue image.
[0042] The tongue image detection method based on the standard color reference of the tongue image detection system based on the standard color reference, wherein the inputting the color features, the texture features, the morphological features and the sublingual collateral features into the constructed multi-modal model for detection and outputting the abnormal detection result of the target tongue image specifically includes:
[0043] The text encoding subunit receives symptom information input by a user through a text encoder, and converts the symptom information into a symptom vector;
[0044] The feature input subunit inputs the color features, the texture features, the morphological features, the sublingual collateral features and the symptom vector into the multi-modal model;
[0045] The encoder in the multi-modal model of the feature fusion subunit utilizes the symptom vector to respectively splice and fuse with the color feature, the texture feature, the shape feature and the sublingual collateral feature, so as to obtain a plurality of modal features;
[0046] The feature output subunit inputs all the modal features into a decoding classifier through the encoder, and the decoding classifier performs prediction on all the modal features to output an abnormality detection result.
[0047] The tongue image detection system based on standard color reference comprises an LED array module, a diffuse polarization module, a light source stabilization module and a color reference module; the LED array module is used for controlling a plurality of LED lamps to cross-illuminate, and converting the illumination light of each LED lamp into corresponding unidirectional soft light through a ring-shaped diffusion cover; the diffuse polarization module is used for converting all the unidirectional soft light into combined soft light, filtering the combined soft light through a plurality of linear polarizers to obtain diffuse light; the light source stabilization module is used for controlling the actual current passing through all the LED lamps in real time, and controlling the camera to delay exposure and shoot under the illumination of the diffuse light to obtain a tongue image; and the color reference module is used for performing multiple color difference calibrations on the tongue image, judging the calibrated image and outputting a detection result. The present application filters the light after diffusing, eliminates the specular reflection, only leaves the diffuse reflection light carrying the real color and texture, and improves the definition of the tongue image. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 is a principle structure diagram of the tongue image detection system based on standard color reference of the present application;
[0049] Figure 2 is another specific principle structure diagram of the tongue image detection system based on standard color reference of the present application;
[0050] Figure 3 is a detector structure diagram of the tongue image detection system based on standard color reference of the present application;
[0051] Figure 4 is a detector structure diagram of the tongue image detection system based on standard color reference of the present application;
[0052] Figure 5 is a tongue image examination and analysis flow chart of the tongue image detection system based on standard color reference of the present application;
[0053] Figure 6 is an image acquisition schematic diagram of the tongue image detection system based on standard color reference of the present application;
[0054] Figure 7is a framework diagram of a tongue appearance detection system based on a standard color reference according to the present application;
[0055] Figure 8 is a processing flow diagram in a preferred embodiment of a tongue appearance detection method based on a standard color reference of a tongue appearance detection system based on a standard color reference according to the present application. DETAILED DESCRIPTION
[0056] In order to make the objectives, technical solutions and advantages of the present application clearer and more apparent, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0057] It should be noted that if the present application embodiments involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement condition, etc. between components in a certain specific posture (as shown in the drawings), and if the specific posture changes, the directional indications also change accordingly.
[0058] In addition, if the present application embodiments involve descriptions of "first", "second", etc., the descriptions of "first", "second", etc. are only for description purposes and should not be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features with "first", "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the realization of a person skilled in the art, and when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0059] One embodiment of the tongue appearance detection system based on a standard color reference according to the preferred embodiment of the present application includes, as shown in Figure 1 The tongue appearance detection system based on a standard color reference includes an LED array module 10, a diffuse polarization module 20, a light source stabilization module 30, and a color reference module 40.
[0060] Among them, the LED array module 10 is used to control the cross-irradiation of multiple LED lights, and converts the irradiation light of each LED light into corresponding unidirectional softened light through an annular diffuser; the diffuse polarization module 20 is used to convert all the unidirectional softened lights into combined softened lights, and filter the combined softened lights through multiple linear polarizers to obtain diffused light; the light source stabilization module 30 is used to control the actual current passing through all the LED lights in real time, and control the camera to delay exposure to obtain the tongue image taken by the camera, wherein the tongue image is taken under the irradiation of the diffuse light; the color reference module 40 is used to perform multiple color difference calibrations on the tongue image, and judge the calibrated image to output the detection results.
[0061] like Figure 2 As shown, another specific embodiment of the tongue image detection system based on standard color reference in the embodiment of the present invention includes: an LED array module 10, a diffuse polarization module 20, a light source stabilization module 30 and a color reference module 40.
[0062] Specifically, the LED array module 10 includes an LED lamp submodule 101 and an annular diffuser 102. The LED lamp submodule 101 includes multiple LED lamps with preset color rendering indices. All of the LED lamps are arranged in a ring on an annular printed circuit board at a preset tilt angle, so that the optical axis of each LED lamp points forward of the central axis of the annular printed circuit board. The annular diffuser 102 is located at a preset distance from all of the LED lamps and covers the front of the LED lamp submodule to soften the light from all of the LED lamps, thereby generating corresponding unidirectional softened light.
[0063] Among them, existing lighting conditions have a huge impact on the color and texture characteristics of tongue images. Natural light is unstable, while parameters such as color temperature and color rendering index of artificial light sources vary, and it is difficult to completely isolate the interference of ambient light. Specular reflection caused by saliva on the tongue surface is also a major interference factor. These factors cause the color of the collected tongue image to be distorted, affecting the accuracy of subsequent analysis. Some systems have tried to use closed or semi-closed collection environments and adopt specific light sources such as halogen lamps or LEDs, but there are still problems with low color temperature or uneven lighting.
[0064] Therefore, the present invention adopts LED array module to solve this problem. Figure 3As shown, the present application adopts a small and compact device to achieve the ultimate portability and the convenience of one-hand operation. The shell is made of lightweight, durable and medical device safety standard-compliant materials, such as medical-grade ABS plastic, polycarbonate or lightweight aluminum alloy. The surface should be smooth and easy to clean. The device is provided with simplified physical buttons, such as power on / off key, image capture key. The status indication can be realized through LED lights (e.g. power status, Bluetooth connection status, in acquisition, acquisition complete, etc.) or small OLED display screen (used to display basic status information or brief results). Since the device will be close to the oral cavity, a replaceable or sterilizable protective cover / guide cover is designed at the front end of the device (the part close to the user's oral cavity) to avoid cross infection. At the same time, the cover can also assist the user to place the tongue at the optimal shooting distance and position.
[0065] Further, for the imaging assembly, a high-resolution miniature CMOS image sensor can be selected, for example, with no less than 5 million effective pixels, to ensure that the fine features of the tongue body (such as tongue fur particles, cracks, and ecchymosis) and the color blocks on the standard color card can be clearly captured. At the same time, a fixed focus or miniature auto-focus lens optimized for close-range (e.g. 5-10 cm) tongue body shooting is equipped. The lens has a suitable field of view to completely accommodate the main part of the extended tongue body and at the same time deploy the standard color reference.
[0066] In view of the low color temperature or uneven illumination, in one embodiment of the present application, a ring-shaped printed circuit board surrounding the camera lens opening is arranged as the base of the LED, and 8 SMD (Surface-Mounted Device) LED chips with high CRI (Color Rendering Index) (high CRI refers to a color rendering index greater than 95) and D65 standard color temperature (color temperature of 6500K) are selected. The high CRI value ensures that the light source can restore the true color of the tongue to the greatest extent. Among them, the 8 LED lights are evenly distributed on the ring-shaped printed circuit board, forming a light ring around the camera lens, providing illumination from multiple directions at the same time, which can greatly offset the shadows caused by the ups and downs of the object. When welding the LED to the PCB, instead of making the light-emitting surface parallel to the PCB (i.e. the light is perpendicular to the front), each LED has an internal inclination angle of 5-10 degrees, and the optical axis points to the center axis of the ring in front. This design forms a "cross fire" type of illumination: the LED located at the top of the ring (12 o'clock position) has a slightly downward inclined light beam, mainly illuminating the middle and lower areas of the tongue surface; the LED located at the bottom of the ring (6 o'clock position) has a slightly upward inclined light beam, mainly illuminating the middle and upper areas of the tongue surface; the same applies to the left and right side LEDs; all light beams form an overlap in the tongue area, which has a remarkable effect on eliminating the fine shadows caused by the tiny convex and concave structures such as the tongue papilla. Compared with the simple ring-shaped lamp with all light beams parallel to the front, this cross layout can make the light better illuminate the grooves of the tongue surface, making the thickness and texture details of the tongue fur more clearly visible.
[0067] Further, in front of the obliquely installed LED array, a ring-shaped diffusion cover is covered. There is a small gap of 2-3 mm between the cover and the LED. When the point light source emitted by the LED lamp passes through the diffusion cover with a certain degree of fogging, the light will be fully scattered, and the scattered light will become very soft, greatly reducing the specular reflection (high light) on the tongue surface caused by saliva. This makes it easier for the subsequent algorithm to analyze the true color and gloss of the tongue, and prevents the dazzling white light spots from misleading the analysis, thereby providing stable and uniform illumination to reduce the influence of environmental light on the quality of the tongue image and ensure that the color reference and tongue are fully illuminated.
[0068] Specifically, the diffuse polarization module 20 comprises a diffuse sub-module 201, a first linear polarizer 202 and a second linear polarizer 203; the diffuse sub-module 201 is used for scattering the single-direction soft light, scattering each single-direction soft light into a corresponding area light source, and obtaining a combined soft light according to all the area light sources; the first linear polarizer 202 is arranged in front of the annular diffuse cover and used for preliminarily filtering the combined soft light to obtain a vertical light in the combined soft light; the second linear polarizer 203 is perpendicular to the first linear polarizer and arranged in front of the camera and used for filtering reflected light to obtain diffuse light; wherein the reflected light represents light reflected after the vertical light irradiates to a target tongue image.
[0069] Wherein, in order to minimize the high light (specular reflection) interference caused by the tongue surface saliva, a diffuse plate can be arranged in front of the illumination system or a cross-polarization technology is adopted. The diffuse plate can make the light more soft and uniform, and reduce direct reflection.
[0070] Wherein, the diffuse plate is usually a semi-transparent material with micro-irregular structure on the surface or inside (such as frosted or milky white polycarbonate plate or acrylic plate), the light is refracted and reflected innumerable times on the internal particles or surface texture, the light reaching the tongue surface is no longer from one or several fixed directions, but from innumerable points on the surface of the diffuse plate, irradiating to the tongue from all directions; since there is no strong incident light with single direction, the condition for forming strong specular reflection is fundamentally destroyed, even if there is reflection, it is the superposition of innumerable weak and blurred reflections, which macroscopically shows uniform and soft brightness, instead of dazzling high light.
[0071] Further, the direction of the irradiation light is filtered by controlling and filtering the vibration direction of the reflected light, the specular reflection light and the diffuse reflection light are selected by the difference in polarization characteristics, and the former is selectively blocked. First, a first linear polarizer (polarizer) is placed in front of the LED light source, the polarizer only allows light with a specific vibration direction (for example, the vertical direction) to pass, at this time, the light irradiating to the tongue surface is all changed into "vertically polarized light". When the vertically polarized light irradiates to the tongue surface, two types of reflected light are generated: (1) specular reflection light (high light), which is direct reflection from the surface of the tongue, which largely retains the polarization characteristics of the original light, so the reflected high light is mainly "vertically polarized light"; (2) diffuse reflection light (effective information), the light enters the tongue fur or the shallow layer of the tongue body, and is reflected after internal scattering. This process will disturb the polarization direction of the light, so that it becomes "unpolarized light" or "partially polarized light" with random vibration direction, and this part of light carries the real color and texture information of the tongue.
[0072] To get effective information, a second linear polarizer (analyzer) is placed in front of the camera lens. The key is to rotate the direction of this polarizer by 90 degrees, so that it only allows "horizontal direction" light to pass through. The "vertical polarization light" from the mirror reflection is almost completely blocked by the second polarizer because the vibration direction does not match. Among the "random polarization light" from the diffuse reflection, there is a component of horizontal vibration, and this part of light can pass through the second polarizer and be captured by the camera. Finally, the image received by the camera almost completely eliminates the mirror reflection (high light) and only leaves the diffuse reflection light carrying the true color and texture. The image will be clearer, and the details such as tongue coating particles and cracks can be more detailed than existing methods.
[0073] Specifically, the light source stabilizing module 30 comprises a current control submodule 301 and a shooting submodule 302; the current control submodule 301 is used for controlling the actual current flowing through all the LED lamps in real time through a feedback control loop, so that the actual current is kept at a preset current value; the shooting submodule 302 is used for lighting all the LED lamps at the actual current within a preset time through a constant current driving chip, and controls the camera to shoot when the PN junction of all the LED lamps reaches an equilibrium point, so as to obtain a tongue image.
[0074] Wherein, when shooting the tongue image, the light source and the color temperature may change, thereby causing unstable illumination of the LED lamp. In the instant of lighting, the temperature of the core semiconductor (PN junction, P-N semiconductor junction) of the LED will rapidly rise. The temperature rise will cause the light efficiency of the LED to decrease (the brightness to decrease) and the color spectrum to deviate (the color temperature to change). Although this process is very fast, its influence cannot be ignored for high-speed and high-precision image acquisition.
[0075] Firstly, the application adopts a dedicated LED constant current driving chip, for example, a DC-DC constant current driving chip (direct current-direct current constant current driving integrated circuit) supporting wide voltage input (such as 2.7V-5.5V) and high output current precision (such as ±1%) is selected. A precise external feedback resistor is used to set the target output current (for example, 150mA). The positive and negative poles of the battery are connected to the input end of the chip, and the LED array is connected to the output end. The constant current driving chip internally integrates a feedback control loop, which monitors the actual current flowing through the LED in real time and high frequency. When the battery voltage decreases or the LED temperature changes, causing the current to deviate from 150mA, the internal switching circuit (such as Boost or Buck) will immediately adjust dynamically, accurately "pull back" the current to the target value of 150mA, and realize real-time monitoring of the circuit.
[0076] Further, in order to eliminate the instability caused by the LED lamp junction temperature drift, the application first controls the constant current driving chip to light up the LED at the target current (150mA) after obtaining the user's click shooting instruction, and lasts for a very short fixed time, for example, 200 milliseconds (ms). The purpose of this short pulse lighting is to quickly warm up the PN junction of the LED and reach a relatively stable "thermal equilibrium" working point (that is, intentionally let the unstable process of LED lamp irradiation occur in advance), when the camera really starts to expose, the temperature and light output of the LED have reached a stable state. Since the preheating time and constant current of each shooting are exactly the same, it is ensured that the lighting conditions of the LED are from a completely same and stable state for each formal shooting, thereby realizing extremely high repeatability.
[0077] Therefore, the application completely eliminates the influence of battery power change on light intensity through constant current driving; through the preheating flash synchronization strategy, the thermal drift stage of the LED start-up is skillfully avoided, and it is ensured that each frame of image collected is completed in a stable and consistent state of the light source; combined with the above two technologies, the system ensures that whether it is continuous snapshot or shooting after several hours, the "sample" of light used for each shooting tends to be consistent in physical characteristics, providing a stable and reliable data basis for subsequent color calibration and diagnostic analysis.
[0078] Specifically, the color reference module 40 includes a global correction submodule 401, a nonlinear calibration submodule 402, and a color judgment submodule 403; the global correction submodule 401 is configured to perform white balance processing on the tongue image by using a first color block to obtain a global calibration image; the nonlinear calibration submodule 402 is configured to perform color calibration on multiple regions of the global calibration image by using a second color block, a third color block, and a fourth color block to obtain a target tongue image; and the color judgment submodule 403 is configured to detect the target tongue image according to the first color block, the third color block, and the fourth color block to obtain a color difference detection result of the target tongue image.
[0079] Among them, although the existing technology adopts external or built-in color cards for color correction (such as ColorChecker), these systems are usually large in size and do not have extreme portability, or the placement and use of the color card rely on manual operation, which is not portable and easy to introduce errors. Therefore, the application seamlessly integrates the standard color reference into a highly portable device, realizing automatic or convenient deployment in the image acquisition process.
[0080] The purpose of the traditional ColorChecker color card is to make the colors of any photo as close to the real world as possible, but the color calibration to be solved by the present application is not aimed at the calibration of different colors, but focuses on the degree of one or more colors (for example, whether the red of the tongue is light red or deep red, etc.), so the micro color card is set to replace other colors in the traditional color card (for example, replace the saturated blue or green color in the traditional color card, which is not present in the tongue image), so that the fuzzy color classification problem is converted into the accurate color comparison problem.
[0081] In another embodiment of the present application, 12 color blocks are arranged on the micro color card, which are divided into four areas:
[0082] Area one is the basic and gray scale reference area, which contains three color blocks, which are used for the most core white balance correction and exposure calibration, including G1: standard white; G2: 50% neutral gray; and G3: standard black.
[0083] Area two is the tongue color gradient area, which contains four color blocks, which are used for the most critical tongue red system reflecting the degree of heat syndrome and the richness and decline of qi and blood, including T1: "light white tongue" reference color: a low saturation color with a slightly pinkish white color, simulating the tongue color when qi and blood are both deficient; T2: "light red tongue" reference color: simulating the healthy tongue color of normal people, which is the "health median line" of all judgments; T3: "red tongue" reference color: a more saturated and bright red color than T2, simulating the tongue color in the early stage of heat syndrome; and T4: "dark red tongue" reference color: a deep red color, even slightly dark red, simulating the severe tongue color of heat invading the blood.
[0084] Area three is the tongue fur color gradient area, which contains three color blocks, which are used to reflect the fur color of pathogenic factors (especially heat evil), including C1: "white fur" reference color: not pure white, but a slightly gray white color simulating normal thin white fur; C2: "light yellow fur" reference color: simulating the thin yellow fur in the early stage of heat syndrome; and C3: "dark yellow fur" reference color: a deep yellow color with higher saturation and slightly brown, simulating the tongue fur color of severe heat evil.
[0085] Area four is the blood stasis index color area, which contains two color blocks, which are used to reflect blood stasis, including S1: "purple tongue" reference color: a blue and dark purple color, commonly seen in cold and blood stasis; and S2: "red purple tongue" reference color: a red and dark purple color, commonly seen in heat and blood stasis.
[0086] Among them, in order to realize integrated setting, the present application sets multiple pop-up mechanisms, which will pop up the micro color card when the tongue image is detected, such as Figure 4For example, for a telescopic / sliding mechanism, a standard color card is fixed on a small sliding arm or tray, which is normally stored inside the device housing. When the user initiates the photo shooting process or presses a specific button, the color card is driven out to a specific position in the camera's field of view (e.g., the edge of the field of view) along a predetermined track by a micro motor or a purely mechanical spring and manual push rod / slider. After the shooting is completed, the mechanism automatically or manually retracts the color card. For a pop-up mechanism, the color card is kept in the storage position by a pre-tightened spring mechanism. When the user triggers the release device (such as a small button or electromagnetic lock), the spring drives the color card to quickly pop up to the preset shooting position. After shooting, it can be pressed back manually or retracted by another mechanism. For a folding / "origami" mechanism, if the color card is made of flexible substrate or is hinged by multiple small hard pieces, it can be designed in a folded form. It is unfolded for deployment and folded for storage to save space. This way requires higher material and precision machinery requirements.
[0087] Further, regardless of the deployment mechanism used, it is necessary to ensure that the position and attitude (angle) of the color card in the camera's field of view each time it is deployed have a high degree of consistency. This is crucial for accurate identification of the color card and its individual color blocks in subsequent image processing. Mechanical limiting, Hall sensor or optical encoder can be used for precise positioning.
[0088] In another embodiment of the present application, a mechanical limiting method is used to achieve precise positioning: first, in the internal structure of the device, a high-precision guide rail or groove is molded by injection molding or CNC machining (Computer Numerical Control Machining), and the sliding bracket of the color card is designed to precisely fit the guide rail, ensuring that it can only move in a linear or predetermined arc; then, at the end of the guide rail, i.e., the ideal position of the fully expanded color card, a physical stopper is designed, which is part of the device shell or internal structure, and its position is fixed. When the sliding bracket moves to this position, it will come into rigid contact with the stopper and cannot move forward; finally, a detent or clasp structure is added at the limiting point to ensure the consistency of the angle (attitude) and prevent rebound or shaking, locking the color card's position from three dimensions and rotational angles, making its attitude highly consistent.
[0089] In another embodiment of the present application, a Hall sensor is used to achieve precise positioning: first, a tiny, strong Neodymium magnet is embedded on a certain point of the color card's sliding bracket, when the user pushes the color card outward, the sliding bracket with the magnet moves, the device's microcontroller (MCU) continuously monitors the output level of the Hall sensor, when the magnet is far away, the sensor output is low (0V). When the sliding bracket reaches the preset ideal position, the magnet on it will move directly above the Hall sensor (or within the effective sensing range), the Hall sensor detects a specific intensity of the magnetic field, its internal circuit will flip, the output level changes from low to high (for example 3.3V) in an instant. When the positioning is completed, the MCU can trigger a feedback mechanism (such as tactile feedback, visual feedback or auditory feedback) to inform the user, achieving non-physical contact, no wear, high precision and smart confirmation function positioning.
[0090] In another embodiment of the present application, an optical encoder is used to achieve precise positioning: first, a transparent or metal film with a special encoding pattern is pasted inside the device along the path of the color card sliding. The incremental encoding strip (uniform black and white stripes) is simpler, and the absolute encoding strip (each millimeter of pattern is a unique binary code) is more accurate, then a tiny optical sensor (containing an infrared LED light source and a photodetector array) is installed on the color card's sliding bracket, facing the encoding strip; further, when the user pushes the color card, the optical reader moves with the sliding bracket, and continuously scans the encoding strip below; the infrared LED light passes through (or reflects) the encoding strip, and the photodetector converts it into a digital position signal according to the on-off or pattern of the received light; then the MCU receives and decodes the data from the optical reader in real time, if an absolute encoder is used, the MCU can instantly get the current position of the color card (for example, "currently at 7.6mm"), at the same time, in the MCU's firmware, a "target position value" is preset, the MCU constantly compares the real-time reading position with the target value, when the position is consistent, the MCU judges that the positioning is completed, and can trigger vibration or light feedback immediately like the Hall sensor scheme, providing micron-level positioning accuracy.
[0091] Further, the global correction submodule 401 comprises a positioning unit 4011, a coefficient calculation unit 4012 and a correction unit 4013; the positioning unit 4011 is configured to determine the accurate area of the neutral gray block in the tongue image according to the first color block; the coefficient calculation unit 4012 is configured to calculate the first color channel gain, the second color channel gain and the third color channel gain in the accurate area; and the correction unit 4013 is configured to update each pixel in the tongue image according to the first color channel gain, the second color channel gain and the third color channel gain, and generate a global calibration map after correction.
[0092] As shown in FIG. 1, after the tongue image captured by the camera is obtained, first, the gray color block in region one is used to complete the white balance and exposure basic correction of the whole image by using the calibration algorithm. Figure 5
[0093] First, the entire color card is positioned in the captured image, and the accurate area of the neutral gray (usually 50% neutral gray) block is further found according to the color card layout; the RGB values of all pixels in this area are read, and the average values of the three channels are calculated, then the gain coefficients of the channels are calculated according to the average values (the gain coefficients help to improve the accuracy of the algorithm output results), finally all pixels in the whole area are traversed and updated, so as to obtain a global calibration map; in order to be more consistent with the visual perception of the human eye or more conducive to subsequent AI analysis, the corrected RGB image can be converted to a device-independent color space.
[0094] For the image after preliminary calibration, the standard color reference needs to be automatically located. The detection can be based on the preset geometric features (such as a specific shape of the frame, corner point) or special positioning markers (such as ArUco markers) on the color card to locate the entire color card. The positioning markers need to meet the following standards: (1) uniqueness: the pattern of the marker must be special enough to rarely appear in the normal shooting environment (including background, user's clothes, skin, etc.); to prevent the algorithm from misidentifying other objects (such as a square pattern on the clothes) in the background as the positioning marker; (2) high contrast: the marker is usually designed with pure black and pure white; the maximum gray difference between black and white enables the algorithm to most clearly separate the outline of the marker during image binarization (i.e., converting the image to black and white), even in poor lighting or slightly blurred images; (3) robustness: the design of the marker should be able to withstand certain degree of perspective transformation (tilt), partial occlusion, motion blur and lighting changes; to ensure that the algorithm can still successfully locate even if the angle is not perfect or the hand has slight shaking when the user holds the device to shoot; (4) high computational efficiency: the algorithm used to detect the marker should be fast enough and not consume too many computing resources; to ensure that the positioning task can be completed smoothly and in real time on mobile devices (such as mobile phones) without causing lag; (5) information carrying capacity: the marker itself should be able to encode some additional information, such as an ID number; to prompt the system about the specific location of the marker (which part of the tongue) so that the corresponding standard color database can be called.
[0095] In another embodiment of the present application, ArUco markers are used: first, a wide and thick black frame and an internal binary coding matrix are used to quickly locate the outline of all squares in the image, and by checking whether the outline is surrounded by a complete black frame, most non-ArUco marker squares (such as windows, chessboards, etc.) are filtered out, achieving efficient filtering in the first step. Then, the verified rectangular region is subjected to perspective transformation to convert it into a standard square, the code of the internal binary matrix is read (for example, white squares are 1 and black squares are 0), a string of binary numbers is obtained, and finally the string of numbers is compared with the preset ArUco "dictionary". If it can be found in the dictionary, the identification is successful, and the ID number of the marker and the precise pixel coordinates of the four corners are returned.
[0096] According to the precise pixel coordinates of the four corners, the precise region can be determined for pixel value positioning, so that color conversion is realized, and finally the target tongue image is obtained. The internal code is used as the positioning code, which can be recognized by the system even in a chaotic background, and through the ID number, multiple different types of calibration boards can be supported; with only ArUco markers, the algorithm can accurately calculate the position and pose (rotation angle) in the three-dimensional space relative to the camera.
[0097] Further, the nonlinear calibration submodule 402 includes a color card positioning unit 4021 and a color calibration unit 4022; the color card positioning unit 4021 is configured to label each of the regions to obtain a corresponding rectangular contour, perform perspective transformation on all the rectangular contours to obtain a corresponding square contour, read a binary matrix code in each of the square contours to obtain a corresponding binary number, and query a preset dictionary according to the binary number to obtain an ID number and a plurality of pixel coordinate points corresponding to each of the square contours; and the color calibration unit 4022 is configured to define a plurality of target coordinate points, input all the pixel coordinate points and all the target coordinate points into a matrix transformation model, output a transformation matrix, reversely calculate all the pixel points in the global calibration graph by using the transformation matrix, obtain the color of each of the pixel points in the global calibration graph, and fill the color of each of the pixel points into a new image to obtain a target tongue image, as shown in Figure 6 .
[0098] When the algorithm accurately finds the four corner points of the color card region in the original image by ArUco markers or other feature detection methods, the pixel coordinates of the four points in the original image are obtained, that is, the source region (an irregular quadrilateral) to be transformed is defined; then an image required by a user (for example, a rectangle) is input; the four corner point coordinates thereof are defined, and the pixel values of the color card are input into the matrix transformation model; the model calculates a transformation matrix by solving a set of linear equations by using the source coordinate points and the target coordinate points; then, by using the transformation matrix, other transformation functions are called to generate a new image with a preset size. For each pixel in the new image, the position of the pixel in the original image (which may not be an integer coordinate) is reversely calculated by using the matrix, and then the color of the point is obtained by interpolation (for example, linear interpolation or cubic interpolation) and filled into the new image. In the new image, each color block is restored to a standard rectangle. The algorithm can now very simply and accurately read the color value of each color block. For example, if the first color block occupies 1 / 6 of the area of the upper left corner, the average value of the colors of all the pixels in the area can be directly obtained, and the colors of the surrounding pixels will not be contaminated. No matter how the device is tilted when the user takes a picture, as long as the four corner points can be recognized, the final generated corrected image is a standard and unified view. This provides high-quality and standardized input data for the subsequent color calibration algorithm.
[0099] Furthermore, the color judgment submodule 403 includes a color difference judgment unit 4031, a feature extraction unit 4032 and a detection unit 4033; the color difference judgment unit 4031 is used to judge the gradient position of the color of all areas of the target tongue image in the first color block, the third color block and the fourth color block. When the color difference detection result meets the preset standard, the target tongue image is segmented to obtain the tongue body area; the feature extraction unit 4032 is used for the feature extraction unit to extract features of the target tongue image according to the tongue body area to obtain the color features, texture features, morphological features and sublingual collateral features of the target tongue image; the detection unit 4033 is used for the detection unit to input the color features, the texture features, the morphological features and the sublingual collateral features into the constructed multimodal model for detection, and output the abnormal detection result of the target tongue image.
[0100] Among them, the detection unit 4033 includes a text encoding subunit 40331, a feature input subunit 40332, a feature fusion subunit 40333 and a feature output subunit 40334; the text encoding subunit 40331 is used to receive the symptom information input by the user through a text encoder and convert the symptom information into a symptom vector; the feature input subunit 40332 is used to input the color feature, the texture feature, the morphological feature, the sublingual venous feature and the symptom vector into a multimodal model; the feature fusion subunit 40333 is used for the encoder in the multimodal model to use the symptom vector to perform splicing and fusion processing with the color feature, the texture feature, the morphological feature and the sublingual venous feature respectively to obtain multiple modal features; the feature output subunit 40334 is used to input all the modal features into a decoding classifier through the encoder, and the decoding classifier predicts all the modal features and outputs an abnormality detection result.
[0101] Further, based on Figure 1 The tongue image detection system based on standard color reference shown in the figure, the tongue image detection method based on standard color reference of the tongue image detection system based on standard color reference of the preferred embodiment of the present invention, as shown in the figure, Figure 7 and Figure 8 As shown, the tongue image detection method based on standard color reference includes the following steps:
[0102] Step S10: The LED array module controls the multiple LED lamps to cross-irradiate, and converts the illumination light of each LED lamp into corresponding unidirectional softened light through the annular diffusion cover.
[0103] Step S20: the diffuse polarization module converts all the unidirectional softened light into combined softened light, and filters the combined softened light through a plurality of linear polarizers to obtain diffused light.
[0104] Step S30, the light source stabilization module controls the actual current passing through all the LED lamps in real time, and controls the camera to delay exposure and take a picture under the irradiation of the diffuse light, to obtain a tongue image.
[0105] Step S40, the color reference module performs multiple color difference calibrations on the tongue image, judges the calibrated image, and outputs a detection result.
[0106] Specifically, the color difference judgment unit judges the gradient positions of the colors of all regions of the target tongue image in the first color block, the third color block and the fourth color block, and when the color difference detection result meets the preset standard, performs segmentation processing on the target tongue image to obtain a tongue body region; the feature extraction unit extracts features of the target tongue image according to the tongue body region, to obtain color features, texture features, morphological features and sublingual collateral features of the target tongue image; and the detection unit inputs the color features, the texture features, the morphological features and the sublingual collateral features into a constructed multi-modal model for detection, and outputs an abnormality detection result of the target tongue image.
[0107] After color calibration, the tongue image still needs further preprocessing to provide clean and standardized input for analysis by the multi-modal large model. First, the tongue body region in the image is accurately separated from the background (such as lips, teeth, oral cavity inner wall and standard color card that has completed its mission), and then visual features related to traditional Chinese medicine diagnosis are extracted from the segmented and color-calibrated tongue body region. These features will be one of the inputs of the multi-modal large model.
[0108] For example, the color features include tongue color: such as pale, pale red, red, purple, purple, etc.; tongue fur color: such as white fur, yellow fur, gray-black fur, etc.; color histogram, color moment (which can form a nine-dimensional feature vector through mean, variance, skewness), principal color analysis, etc. can be used to quantify color. Texture features include thickness, dryness, rotting, peeling, etc. of tongue fur; texture of tongue body, such as cracks; gray level co-occurrence matrix, local binary pattern, wavelet transform, etc. can be used to extract texture features. Morphological features include tongue size and shape: such as thick tongue, thin tongue, tooth mark tongue, crack tongue, prickle tongue, ecchymosis tongue, etc., which can calculate the area, perimeter, aspect ratio, symmetry, edge smoothness, tooth mark number and depth, crack distribution and shape, etc. of the tongue body. Sublingual collateral features include the color, shape, thickness, tortuosity, and presence or absence of ecchymosis points of the two sublingual veins, which usually need the user to roll the tongue up to take a picture. Other features also include the amount of tongue surface fluid (reflecting dryness) and the like.
[0109] Specifically, the text encoding subunit receives symptom information input by a user through a text encoder and converts the symptom information into a symptom vector; the feature input subunit inputs the color feature, the texture feature, the shape feature, the sublingual collateral feature, and the symptom vector into a multi-modal model; the feature fusion subunit performs splicing and fusion processing on the symptom vector and the color feature, the texture feature, the shape feature, and the sublingual collateral feature respectively by using an encoder in the multi-modal model, to obtain a plurality of modal features; and the feature output subunit inputs all the modal features into a decoding classifier through the encoder, and the decoding classifier performs prediction on all the modal features and outputs an abnormality detection result.
[0110] Among them, the multi-modal model receives the tongue appearance visual features extracted by the pre-processing module (or directly receives the calibrated tongue image block) as input, and simultaneously accesses the text information (mainly including the current symptoms, brief medical history, and living habits input by the user, which can be converted into vector representation through a text encoder) input by the user APP, and then the multi-modal model is responsible for effectively fusing the tongue appearance features from the visual encoder and the text features from the text encoder, and then analyzing by using the decoder: if the target is to judge which pre-defined TCM syndrome type (such as spleen deficiency, liver fire, etc.) or health status, the decoder can be a series of fully connected layers followed by a Softmax activation function; if the target is to generate a descriptive diagnostic report (for example, “tongue pale red, thin white moss, edge with tooth marks, consider as spleen deficiency syndrome…”), the decoder can be a decoding part of a large language model, which generates a text report autoregressively conditioned on the multi-modal features; at the same time, it can also evaluate the state of the corresponding viscera according to the tongue surface partition (such as the tongue tip corresponding to the heart and lungs, the tongue middle corresponding to the spleen and stomach, the tongue root corresponding to the kidney, and the tongue sides corresponding to the liver and gallbladder) respectively.
[0111] The application provides a tongue image detection system and method based on standard color reference, which comprises an LED array module, a diffuse polarization module, a light source stabilization module and a color reference module; the LED array module is used for controlling multiple LED lamps to cross-illuminate and converting the illumination light of each LED lamp into corresponding unidirectional soft light through a ring-shaped diffusion cover; the diffuse polarization module is used for converting all the unidirectional soft light into combined soft light, filtering the combined soft light through multiple linear polarizers to obtain diffuse light; the light source stabilization module is used for controlling the actual current passing through all the LED lamps in real time and controlling the camera to delay exposure and take pictures under the illumination of the diffuse light to obtain a tongue image; and the color reference module is used for performing multiple color difference calibrations on the tongue image and judging the calibrated image to output a detection result. The application filters the light after diffusing the light, eliminates the specular reflection, only leaves the diffuse reflection light carrying the real color and texture, and improves the definition of the tongue image.
[0112] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0113] It should be understood that the application is not limited to the above examples, and can be improved or changed by those skilled in the art according to the above description, and all these improvements and changes shall fall within the protection scope of the appended claims of the application.
Claims
1. A tongue image detection system based on standard color reference, characterized by, The tongue image detection system based on the standard color reference comprises an LED array module, a diffuse polarization module, a light source stabilization module and a color reference module. The LED array module is used for controlling multiple LED lamps to cross-illuminate and converting the illumination light of each LED lamp into corresponding single-direction softened light through a ring-shaped diffusion cover. The diffuse polarization module is used for converting all the single-direction softened light into combined softened light, filtering the combined softened light through multiple linear polarizers to obtain diffuse light. The light source stabilization module is used for controlling the actual current flowing through all the LED lamps in real time and controlling the camera to delay exposure and take a picture under the illumination of the diffuse light to obtain a tongue image. The color reference module is used for performing multiple color difference calibrations on the tongue image and judging the calibrated image to output a detection result.
2. The standard color reference based tongue inspection system according to claim 1, characterized in that, The LED array module comprises an LED lamp sub-module and a ring-shaped diffusion cover. The LED lamp sub-module comprises multiple LED lamps with preset color rendering indexes, and all the LED lamps are arranged in a ring shape on a ring-shaped printed circuit board at a preset inclination angle, so that the optical axis of each LED lamp points to the front of the central axis of the ring-shaped printed circuit board. The ring-shaped diffusion cover covers the front of the LED lamp sub-module at a preset distance from all the LED lamps to soften the illumination light of all the LED lamps to obtain corresponding single-direction softened light.
3. The standard color reference based tongue inspection system according to claim 1, characterized in that, The diffuse polarization module comprises a diffuse sub-module, a first linear polarizer and a second linear polarizer. The diffuse sub-module is used for scattering the single-direction softened light to obtain corresponding area light sources, and combined softened light is obtained according to all the area light sources. The first linear polarizer is arranged in front of the ring-shaped diffusion cover and is used for preliminarily filtering the combined softened light to obtain vertical light in the combined softened light. The second linear polarizer is perpendicular to the first linear polarizer and is arranged in front of the camera and is used for filtering reflected light to obtain diffuse light. The reflected light represents the light reflected by the target tongue image after being illuminated by the vertical light.
4. The standard color reference based tongue inspection system according to claim 1, characterized in that, The light source stabilization module comprises a current control sub-module and a shooting sub-module. The current control sub-module is used for controlling the actual current flowing through all the LED lamps in real time through a feedback control loop to keep the actual current at a preset current value. The shooting sub-module is used for turning on all the LED lamps at the actual current for a preset time through a constant current driving chip, controlling the camera to take a picture when the PN junction of all the LED lamps reaches an equilibrium point to obtain a tongue image.
5. The standard color reference based tongue inspection system according to claim 1, wherein The color reference module comprises a global correction sub-module, a non-linear calibration sub-module and a color judgment sub-module. The global correction sub-module is used for performing white balance processing on the tongue image through a pre-set first color block to obtain a global calibration image. The nonlinear calibration submodule is configured to perform color calibration on multiple regions of the global calibration map respectively by using a pre-set second color block, a third color block, and a fourth color block, to obtain a target tongue image; The color judgment submodule is configured to perform detection on the target tongue image according to the first color block, the third color block, and the fourth color block respectively, to obtain a color difference detection result of the target tongue image.
6. The standard color reference based tongue inspection system according to claim 5, characterized in that, The global correction submodule includes a positioning unit, a coefficient calculation unit, and a correction unit; The positioning unit is configured to determine an accurate region of a neutral gray block in the tongue image according to the first color block; The coefficient calculation unit is configured to calculate a first color channel gain, a second color channel gain, and a third color channel gain in the accurate region; The correction unit is configured to perform traversal update on each pixel in the tongue image according to the first color channel gain, the second color channel gain, and the third color channel gain, to generate a corrected global calibration map.
7. The standard color reference based tongue inspection system according to claim 5, characterized in that, The nonlinear calibration submodule includes a color card positioning unit and a color calibration unit; The color card positioning unit is configured to mark each region to obtain a corresponding rectangular contour, perform perspective transformation on all the rectangular contours to obtain a corresponding square contour, read a binary matrix code in each square contour to obtain a corresponding binary number, and query the binary number in a pre-set dictionary to obtain an ID number and multiple pixel coordinate points corresponding to each square contour; The color calibration unit is configured to define multiple target coordinate points, input all the pixel coordinate points and all the target coordinate points into a matrix transformation model, output a transformation matrix, inversely calculate all the pixel points in the global calibration map by using the transformation matrix, obtain the color of each pixel point in the global calibration map, and fill the color of each pixel point into a new image to obtain a target tongue image.
8. A standard color reference-based tongue image detection method based on the standard color reference-based tongue image detection system according to any one of claims 1 to 7, characterized by, The tongue image detection method based on a standard color reference includes: The LED array module controls multiple LED lamps to cross-illuminate, and converts the illumination light of each LED lamp into corresponding unidirectional soft light through a ring-shaped diffusion cover; The diffuse polarization module converts all the unidirectional soft light into combined soft light, filters the combined soft light through multiple linear polarizing plates, and obtains diffuse light; The light source stabilization module controls the actual current passing through all the LED lamps in real time, controls the camera to delay exposure, and captures an image under the illumination of the diffuse light to obtain a tongue image; The color reference module performs multiple color difference calibrations on the tongue image, judges the calibrated image, and outputs a detection result.
9. The standard color reference based tongue inspection method according to claim 8, characterized in that, The judgment on the calibrated image and the output of the detection result specifically include: A color difference judgment unit judges the gradient positions of the colors of all regions of a target tongue image in a first color block, a third color block, and a fourth color block, performs segmentation processing on the target tongue image when the color difference detection result meets a pre-set standard, and obtains a tongue body region. The feature extraction unit extracts features of the target tongue image according to the tongue body region, and obtains color features, texture features, morphological features and sublingual collateral features of the target tongue image; The detection unit inputs the color features, the texture features, the morphological features and the sublingual collateral features into the constructed multi-modal model for detection, and outputs an abnormality detection result of the target tongue image.
10. The standard color reference based tongue inspection method according to claim 9, characterized in that, The color features, the texture features, the morphological features and the sublingual collateral features are input into the constructed multi-modal model for detection, and an abnormality detection result of the target tongue image is output, specifically including: The text encoding subunit receives symptom information input by a user through a text encoder, and converts the symptom information into a symptom vector; The feature input subunit inputs the color features, the texture features, the morphological features, the sublingual collateral features and the symptom vector into the multi-modal model; The feature fusion subunit utilizes the symptom vector to respectively splice and fuse the color features, the texture features, the morphological features and the sublingual collateral features in the encoder of the multi-modal model, and obtains a plurality of modal features; The feature output subunit inputs all the modal features into a decoding classifier through the encoder, the decoding classifier predicts all the modal features, and outputs an abnormality detection result.