Face multi-modal information acquisition equipment

By integrating multimodal information acquisition equipment, the problem of lack of objectivity in the diagnosis of orbital disease is solved, efficient and synchronous data acquisition is achieved, the accuracy and reliability of the diagnosis are improved, and personalized treatment is supported.

CN120391998APending Publication Date: 2025-08-01SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234542.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing diagnosis methods for orbital diseases lack objective and comprehensive data support, which makes the diagnosis highly subjective and makes it difficult to conduct accurate diagnosis and treatment evaluation.

Method used

Design a multimodal facial information acquisition device, integrating RGBD camera system, eye tracking system, lighting system, audio input system and microprocessing system, which can synchronize RGBD images, videos, eye movement data, three-dimensional models and audio data of the patient's face, providing rich diagnostic data.

Benefits of technology

It significantly improves the accuracy and reliability of the diagnosis, reduces patient fatigue and discomfort, ensures consistency and comprehensiveness of the data, provides doctors with detailed facial and eye data, and supports the formulation of personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120391998A_ABST
    Figure CN120391998A_ABST
Patent Text Reader

Abstract

The invention discloses a face multi-modal information acquisition device, which comprises a hemispherical shell with an opening facing the face of a patient; the head of the patient is located at the sphere center of the hemispherical shell; the RGBD camera system is fixed to the inner side face of the hemispherical shell and used for shooting RGBD images and videos of the face of the patient; the eyeball tracking system is fixed to the inner side face of the hemispherical shell and used for tracking the human eye rotation angle and the fixation line direction of the patient; the illumination system is fixed on the inner side surface of the hemispherical shell and is used for providing required brightness when RGBD images and videos are shot and providing eyeball observation orientation guidance for a patient; the audio input system is fixed to the inner side face of the hemispherical shell and used for collecting audio data of doctors and patients; and the micro-processing system is in communication connection with the RGBD camera system and is used for measuring and analyzing the RGBD images and videos so as to generate a three-dimensional model of the face of the patient. According to the invention, integrated, efficient and synchronous multi-modal information acquisition is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical devices, and particularly to a facial multi-modal information acquisition device for the diagnosis and treatment of orbital diseases. Background Art

[0002] Orbital diseases refer to various diseases occurring in the orbit, including orbital tumors, inflammations, traumas, congenital deformities, etc.; common orbital diseases include thyroid eye disease (TED), orbital hemangioma, orbital fracture, orbital inflammation, etc. These diseases can cause symptoms such as proptosis, eye pain, vision loss, diplopia, etc., and may even lead to blindness in severe cases.

[0003] Currently, the diagnosis of orbital diseases mainly relies on the clinical experience of doctors; doctors judge the condition through the patient's medical history, symptom manifestations, physical examinations or by means of imaging examinations such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), etc. However, this diagnostic method has certain subjectivity because the experience and judgment criteria of different doctors vary. Therefore, traditional examination methods lack objective and comprehensive data support and are difficult to perform accurate diagnosis and treatment evaluation. Summary of the Invention

[0004] The purpose of the present invention is to provide a facial multi-modal information acquisition device that can simultaneously acquire five-modal information of the patient's facial RGBD images, videos, eye movement data, three-dimensional models of the face, and audio data, greatly improving the acquisition efficiency and ensuring the consistency and accuracy of the data, thereby significantly improving the accuracy and reliability of diagnosis.

[0005] To achieve the above purpose, the present invention is realized through the following technical solutions:

[0006] A facial multi-modal information acquisition device for the diagnosis and treatment of orbital diseases in patients, comprising:

[0007] A hemispherical housing; the opening of the hemispherical housing faces the patient's face, and the patient's head is located at the center of the sphere of the hemispherical housing;

[0008] An RGBD camera system, fixed on the inner side surface of the hemispherical housing, for taking RGBD images of the patient's face and taking videos; each pixel point in the RGBD image includes R, G, B color information and depth information;

[0009] An eye tracking system, fixed on the inner side surface of the hemispherical housing, for tracking and obtaining the rotation angle of the patient's human eyes and the gaze line direction;

[0010] A lighting system, fixed to the inner side surface of the hemispherical housing, is used to provide the required brightness when taking RGBD images and videos and to provide guidance on the eye observation orientation for the patient;

[0011] An audio input system, fixed to the inner side surface of the hemispherical housing, is used to collect audio data of the doctor and the patient; and

[0012] A microprocessing system, communicatively connected to the RGBD camera system, is used to measure and analyze the RGBD images and the videos to generate a three-dimensional model of the patient's face.

[0013] Optionally, the hemispherical housing includes a first semi-circular arc and a second semi-circular arc, and the plane where the first semi-circular arc is located is a horizontal plane, and the plane where the second semi-circular arc is located is a vertical plane perpendicular to the horizontal plane; the centers of both the first semi-circular arc and the second semi-circular arc are the center of the sphere of the hemispherical housing; the first semi-circular arc is tangent to the second semi-circular arc, and the tangent point is the midpoint of both;

[0014] The RGBD camera system includes: nine RGBD cameras; the nine RGBD cameras are respectively denoted as the first RGBD camera, the second RGBD camera, the third RGBD camera, the fourth RGBD camera, the fifth RGBD camera, the sixth RGBD camera, the seventh RGBD camera, the eighth RGBD camera, and the ninth RGBD camera;

[0015] Among them, the first RGBD camera is located at the tangent point, the second to fifth RGBD cameras are sequentially arranged on the first semi-circular arc, and the sixth to ninth RGBD cameras are sequentially arranged on the second semi-circular arc; and the second RGBD camera and the fifth RGBD camera are respectively located at the two end points of the first semi-circular arc, the third RGBD camera is located between the second RGBD camera and the first RGBD camera, and the fourth RGBD camera is located between the first RGBD camera and the fifth RGBD camera; the sixth RGBD camera and the seventh RGBD camera are located between the first end point of the second semi-circular arc and the first RGBD camera, and the seventh RGBD camera is close to the first RGBD camera; the eighth RGBD camera and the ninth RGBD camera are located between the first RGBD camera and the second end point of the second semi-circular arc, and the eighth RGBD camera is close to the first RGBD camera.

[0016] Optionally, the sixth RGBD camera, the center of the sphere of the hemispherical housing, and the ninth RGBD camera form a first angle, the vertex of the first angle is the center of the sphere of the hemispherical housing, and the degree of the first angle is 110 degrees to 130 degrees.

[0017] Optionally, the eye tracking system includes:

[0018] An eye movement camera, fixed on the second semi-circular arc, between the first RGBD camera and the eighth RGBD camera;

[0019] An eye tracking display, fixed on the second semi-circular arc, between the first RGBD camera and the eye movement camera.

[0020] Optionally, a second angle is formed by the first RGBD camera, the center of the spherical shell, and the eye movement camera. The vertex of the second angle is the center of the spherical shell, and the degree of the second angle is 20 degrees to 40 degrees.

[0021] Optionally, the lighting system includes:

[0022] A first lighting source and a second lighting source, located above the plane of the first semi-circular arc, fixed on the inner side of the hemispherical shell, for providing the required brightness when shooting RGBD images and videos; and the first lighting source and the second lighting source are respectively located on both sides of the plane of the second semi-circular arc;

[0023] Nine LED lights, evenly arranged in 3 rows and 3 columns on the inner side of the hemispherical shell, for guiding the patient to observe the eyes.

[0024] Optionally, the audio input system includes at least one microphone; the microphone is fixed on the inner side of the hemispherical shell and is at the same horizontal height as the eye tracking display.

[0025] Optionally, a headrest for supporting the patient's head is provided at the center of the spherical shell of the hemispherical shell.

[0026] Optionally, the hemispherical shell has a contracted state and an expanded state; when diagnosing and treating orbital diseases of the patient, the hemispherical shell is in the expanded state; when not diagnosing and treating orbital diseases of the patient, the hemispherical shell is in the contracted state; and the hemispherical shell is made of light-tight material.

[0027] Optionally, the facial multi-modal information acquisition device further includes:

[0028] A fixed platform, fixedly connected to the hemispherical shell in the expanded state, for carrying the hemispherical shell;

[0029] A lifting bracket, fixedly connected to the fixed platform, for supporting the fixed platform and driving the fixed platform to reciprocate in the vertical direction.

[0030] The present invention has at least one of the following advantages:

[0031] The facial multi-modal information acquisition device provided by the present invention integrates an RGBD camera system, an eye tracking system, an illumination system, an audio input system, and a microprocessing system. It can simultaneously acquire five types of information, namely RGBD images, videos, eye movement data, 3D models of the face, and audio data of the patient's face during one operation process, providing rich diagnostic data for doctors to help identify key symptoms such as abnormal orbital morphology and eye movement disorders, and significantly improving the accuracy and reliability of diagnosis. Compared with the prior art, the present invention realizes integrated, efficient, and synchronous multi-modal information acquisition, significantly improving the acquisition efficiency and reducing the patient's fatigue and discomfort caused by multiple shootings. At the same time, the present invention can acquire multi-modal information at the same time and in the same environment, ensuring the consistency and accuracy of the data and avoiding problems such as inconsistent shooting scenarios and patient states caused by scattered devices. The present invention not only improves the patient's cooperation but also reduces the doctor's operation time and workload, ensuring the comprehensiveness and accuracy of multi-modal data.

[0032] The present invention includes components such as 9 RGBD cameras, an eye movement camera, 9 LED lights, and an illumination light source, and the set of components is relatively comprehensive. The 9 RGBD cameras can achieve multi-angle coverage and comprehensively photograph the patient's face to obtain detailed facial information from different angles; the eye movement camera can accurately track the patient's eye movement conditions and provide dynamic eye movement data; the 9 LED lights are used to indicate the deflection angle of the patient's eyes to ensure the accuracy and consistency of data acquisition; the illumination light source is used to keep the acquisition background bright. Through the combination of these technologies, the present invention can provide more comprehensive and detailed facial and eye data to help doctors make more accurate diagnoses and treatment evaluations.

[0033] The present invention can also generate a detailed 3D model of the patient's face, which not only helps doctors comprehensively observe the changes in the orbital structure but also provides comparison data before and after surgery, providing important support for the formulation of treatment plans and the evaluation of treatment effects.

[0034] The multi-modal information acquisition function of the present invention provides a solid foundation for the artificial intelligence automatic diagnosis of orbital diseases. The rich and comprehensive multi-modal data enables artificial intelligence algorithms to better learn and identify the characteristics of orbital diseases, automatically generate diagnostic results, thereby improving the diagnostic efficiency and reducing the doctor's workload. At the same time, the automatic diagnosis system can continuously improve the accuracy and reliability of diagnosis through continuous learning and optimization.

[0035] The present invention is easy to operate, and the data acquisition process is fast and efficient, reducing the patient's discomfort. Doctors can more comprehensively and timely understand the patient's condition through the intuitive data and images provided by the present invention and formulate personalized treatment plans. The audio records the patient's subjective symptom descriptions, further enriching the diagnostic information and helping doctors conduct comprehensive evaluations. Brief Description of the Drawings

[0036] Figure 1 It is a schematic structural diagram of a facial multi-modal information acquisition device provided by an embodiment of the present invention;

[0037] Figure 2 It is a schematic distribution diagram of an illumination system and an audio input system in a facial multi-modal information acquisition device provided by an embodiment of the present invention;

[0038] Figure 3 It is an orientation diagram when a facial multi-modal information acquisition device provided by an embodiment of the present invention acquires an image;

[0039] Figure 4 It is a topology diagram of a data synchronization structure in a facial multi-modal information acquisition device provided by an embodiment of the present invention;

[0040] Figure 5 It is a data processing flow chart of a microprocessing system in a facial multi-modal information acquisition device provided by an embodiment of the present invention. Detailed implementation manners

[0041] The following further elaborates in detail on a facial multi-modal information acquisition device proposed by the present invention in conjunction with the accompanying drawings and specific implementation manners. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are in a very simplified form and all use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention. In order to make the purpose, features, and advantages of the present invention more obvious and understandable, please refer to the accompanying drawings. It should be known that the structures, scales, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed by the present invention.

[0042] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0043] As described in the background art, traditional methods for diagnosing orbital diseases lack objective and comprehensive data support. This indicates that for accurate diagnosis and treatment evaluation of orbital diseases, multiple modalities of information are often required, but the acquisition of this information is usually scattered and needs to be collected separately by multiple devices. For example, a doctor needs to use one device to take 2D photos, another device to record videos, a third device to perform 3D reconstruction, a fourth device to track eye movement data, and finally an additional device to record audio. These devices cannot collect multiple modalities of information synchronously and in the same place. This is not only time-consuming and laborious, but also the switching between different devices and the adjustment of the patient's posture will result in inconsistent shooting scenarios and patient states, leading to a mismatch in the timing between different devices and affecting the accuracy and consistency of the data. In addition, patients are prone to fatigue and discomfort during multiple adjustments and repositionings, resulting in a decrease in their cooperation, which in turn affects the quality and consistency of data collection. The acquisition environment and parameters of each device may also vary, which will further exacerbate the inconsistency of the data and the error of the diagnosis result. Overall, this method of decentralized collection by multiple devices is inefficient, cumbersome to operate, and prone to various potential problems, seriously affecting the accuracy and reliability of the diagnosis. As can be seen from the above, the method of decentralized collection by multiple devices has great limitations in the diagnosis and treatment of orbital diseases, cannot provide comprehensive and accurate multi-modal data support, and seriously affects the accuracy of the diagnosis and the decision-making of the treatment plan.

[0044] Combined with the attached Figures 1 - 2As shown in the figure, this embodiment provides a facial multi-modal information acquisition device for diagnosing and treating orbital diseases of patients. The facial multi-modal information acquisition device includes: a hemispherical housing 110, an RGBD camera system, an eye tracking system, an illumination system, an audio input system, and a microprocessing system. The opening 1101 of the hemispherical housing 110 faces the face of the patient, and the head of the patient is located at the center O of the sphere of the hemispherical housing 110. The RGBD camera system is fixed on the inner side of the hemispherical housing 110 for taking RGBD images and videos of the face of the patient; and each pixel point in the RGBD image includes R, G, B color information and depth information. The eye tracking system is fixed on the inner side of the hemispherical housing 110 for tracking and obtaining the rotation angle of the human eye and the gaze line direction (i.e., eye movement data) of the patient. The illumination system is fixed on the inner side of the hemispherical housing 110 for providing the required brightness when taking RGBD images and videos and for guiding the patient to observe the position of the eyeball. The audio input system is fixed on the inner side of the hemispherical housing 110 for collecting audio data of doctors and patients. The microprocessing system is communicatively connected to the RGBD camera system for measuring and analyzing the RGBD images and the videos to generate a three-dimensional model of the face of the patient.

[0045] The facial multi-modal information acquisition device provided in this embodiment specifically for orbital diseases integrates an RGBD camera system, an eye tracking system, an illumination system, an audio input system, and a microprocessing system, enabling the facial multi-modal information acquisition device to efficiently and integrally synchronously acquire five types of facial modal information (i.e., RGBD images, videos, eye movement data, three-dimensional models of the face, and audio data of the patient's face). The facial multi-modal information acquisition device provided in this embodiment effectively improves the acquisition efficiency of facial modal information and the cooperation degree of patients, and also ensures the consistency and accuracy of data, which can help doctors make more accurate diagnoses and treatment evaluations, and at the same time provides a solid foundation for the artificial intelligence automatic diagnosis of orbital diseases.

[0046] Specifically, in the facial multi-modal information acquisition device provided in this embodiment, the hemispherical housing 110 has a contracted state and an expanded state; among them, when diagnosing and treating orbital diseases of the patient, the hemispherical housing is in the expanded state (such as Figure 1 and Figure 2As shown; when the patient is not undergoing orbital disease diagnosis and treatment, the hemispherical housing is in a contracted state, that is, when the facial multi-modal information acquisition device is not in use, it is contracted to facilitate the storage of the facial multi-modal information acquisition device and make its occupied space smaller. Optionally, the hemispherical housing 110 can adopt a contraction and expansion structure similar to that of an umbrella, which is not limited here. Optionally, the hemispherical housing 110 is made of a light-tight material to block ambient light, thereby avoiding adverse effects of ambient light on the shooting of RGBD images and videos. Optionally, a headrest for supporting the patient's head (not shown in the figure) is provided at the center of the sphere O of the hemispherical housing 110 to support the patient's head when shooting RGBD images and videos, so that the patient keeps the head and face fixed, thereby improving the shooting accuracy and efficiency of RGBD images and videos, and at the same time saving the patient's physical strength, but the present invention is not limited thereto.

[0047] Please continue to refer to Figure 1 , the hemispherical housing 110 includes a first semi-circular arc and a second semi-circular arc, and the plane where the first semi-circular arc D⌒AE is located is a horizontal plane, and the plane where the second semi-circular arc S⌒AT is located is a vertical plane perpendicular to the horizontal plane; the centers of the first semi-circular arc D⌒AE and the second semi-circular arc S⌒AT are both the center of the sphere O of the hemispherical housing 110; the first semi-circular arc D⌒AE is tangent to the second semi-circular arc S⌒AT, and the tangent point is the midpoint of both.

[0048] The RGBD camera system includes: nine RGBD cameras, each of which is composed of a structured light camera and an RGB camera. The nine RGBD cameras are respectively denoted as the first RGBD camera A, the second RGBD camera D, the third RGBD camera F, the fourth RGBD camera G, the fifth RGBD camera E, the sixth RGBD camera B, the seventh RGBD camera J, the eighth RGBD camera I, and the ninth RGBD camera C. Among them, the first RGBD camera A is located at the tangent point, the second to fifth RGBD cameras are sequentially arranged on the first semi-circular arc D⌒AE, and the sixth to ninth RGBD cameras are sequentially arranged on the second semi-circular arc S⌒AT. Moreover, the second RGBD camera D and the fifth RGBD camera E are respectively located at the two end points of the first semi-circular arc D⌒AE, the third RGBD camera F is located between the second RGBD camera D and the first RGBD camera A, and the fourth RGBD camera G is located between the first RGBD camera A and the fifth RGBD camera E. The sixth RGBD camera B and the seventh RGBD camera J are located between the first end point S of the second semi-circular arc S⌒AT and the first RGBD camera A, and the seventh RGBD camera J is close to the first RGBD camera A. The eighth RGBD camera I and the ninth RGBD camera C are located between the first RGBD camera A and the second end point T of the second semi-circular arc S⌒AT, and the eighth RGBD camera I is close to the first RGBD camera A.

[0049] Specifically, in one embodiment, the seventh RGBD camera J and the eighth RGBD camera I are backup cameras, and the first to sixth RGBD cameras and the ninth RGBD camera are main cameras. Optionally, the sixth RGBD camera B, the center O of the hemispherical housing 110, and the ninth RGBD camera C form a first angle ∠BOC. The vertex of the first angle ∠BOC is the center O of the hemispherical housing 110, and the degree of the first angle ∠BOC is 110 degrees to 130 degrees, so that the nine RGBD cameras can achieve multi-angle coverage and comprehensively photograph the patient's face to obtain detailed facial information from different angles. Preferably, the degree of the first angle ∠BOC is 120 degrees, but the present invention is not limited thereto.

[0050] More specifically, the RGBD images captured by the RGBD camera system can clearly display the facial images of the patient. The captured RGBD images can provide static facial features of the patient, such as eyelid morphology, skin condition, and facial symmetry, which can help doctors conduct preliminary screening and assessment and identify surface symptoms such as ptosis and eye swelling. The video captured by the RGBD camera system can completely and clearly display the dynamic eye information of the patient. The captured video can capture dynamic changes and eye movements, provide behavioral and functional information, and help doctors detect dynamic diseases such as eye movement disorders, incomplete eyelid closure, and nystagmus, and provide continuous time-series data for analysis.

[0051] Please continue to refer to Figure 2 , the lighting system includes: a first lighting source L1, a second lighting source L2, and nine LED lights (LL1 to LL9); the first lighting source L1 and the second lighting source L2 are located above the plane of the first semi-circular arc D⌒AE and are fixed on the inner side surface of the hemispherical housing 110, and are used to provide the required brightness when capturing RGBD images and videos to keep the acquisition background bright; and the first lighting source L1 and the second lighting source L2 are respectively located on both sides of the plane of the second semi-circular arc S⌒AT. The nine LED lights (LL1 to LL9) are evenly arranged in three rows and three columns on the inner side surface of the hemispherical housing 110, and are used to provide eye observation direction guidance to the patient, so as to facilitate the RGBD camera system to capture corresponding RGBD images, thereby improving the acquisition efficiency of the RGBD images of the patient.

[0052] Specifically, as Figure 2 and Figure 3 shown, the eye observation directions include: frontal, up, down, right, upper right, lower right, left, upper left, and lower left, that is, the eye observation directions are set in one-to-one correspondence with the nine LED lights; for example, when the LED light LL2 corresponding to "up" lights up, the patient's eyes look up, and at this time the RGBD camera system can capture the RGBD image of the patient's eyes looking up. More specifically, for the convenience of doctors to conduct accurate diagnosis and treatment, an RGBD image needs to be captured in each eye observation direction. In this case, RGBD images of nine eye observation directions can be obtained. When capturing the three directions of lower right, down, and lower left, the eyelids need to be pulled open so that the RGBD camera system can better capture the eyes. Optionally, the light emitted by the LED lights is red light, but the present invention is not limited thereto.

[0053] Please continue to refer to Figure 1 and Figure 2, the eye tracking system includes: an eye movement camera H, fixed on the second semi-circular arc S⌒AT, between the first RGBD camera A and the eighth RGBD camera I; and an eye tracking display K1, fixed on the second semi-circular arc S⌒AT, between the first RGBD camera A and the eye movement camera H. Optionally, the eye tracking display K1 is disposed close to the eye movement camera H, and the eye tracking display K1 will display fixation information, etc., to accurately track the eye movement of the patient through the cooperation of the eye movement camera H and the eye tracking display K1, so as to obtain the human eye rotation angle and the fixation line direction (i.e., eye movement data) of the patient.

[0054] Specifically, in one embodiment, the first RGBD camera A, the center O of the hemispherical housing 110, and the eye movement camera H form a second angle ∠AOH. The vertex of the second angle ∠AOH is the center O of the hemispherical housing 110, and the degree of the second angle ∠AOH is 20 degrees to 40 degrees, so that the eye movement camera H can comprehensively track the eye movement state of the patient. Preferably, the degree of the second angle ∠AOH is 30 degrees, but the present invention is not limited thereto.

[0055] More specifically, the eye movement data obtained by the eye tracking system records the eye movement trajectory and speed; according to the eye movement data, the movement pattern and reaction of the eye can be analyzed, which reflects the functional state of the nervous system and the eye muscles, and helps the doctor identify nervous system diseases and eye muscle dysfunction.

[0056] Please continue to refer to Figure 2 , the audio input system includes at least one microphone M1; the microphone M1 is fixed on the inner side of the hemispherical housing 110 and is at the same horizontal height as the eye tracking display K1, so as to facilitate the simultaneous collection of the audio data of the doctor and the patient, so that when the doctor interviews the patient, the description of the patient's subjective feelings such as eye discomfort and pain can be recorded, and then the symptom information that cannot be captured in the image data can be supplemented, helping the doctor to more comprehensively understand the patient's symptoms and conduct comprehensive evaluation and personalized diagnosis.

[0057] In addition, in this embodiment, the facial multi-modal information acquisition device further includes: a display system, which consists of a desktop display and corresponding GUI software, and is used to prompt the patient to collect information and guide the doctor to operate. The microprocessing system can be composed of a computing PC processor, computing software, and a PTP network synchronization control component, providing computing power support and a software operation platform, and controlling the data acquisition synchronization trigger of the RGBD camera system, the eye tracking system, the display system, the lighting system, and the audio input system.

[0058] Specifically, asFigure 5 As shown, the microprocessing system can achieve single-angle output and stitching output of RGBD images (depth and point cloud images) and videos; and achieve the output of a three-dimensional model (i.e., 3D model) of the patient's face. The output three-dimensional model includes texture, color, feature points, and three-dimensional facial structure information, which can display the three-dimensional morphology of the orbit and face in detail, and is suitable for evaluating structural diseases such as orbital tumors and orbital fractures, and assisting in formulating surgical plans and evaluating the effects. More specifically, the microprocessing system can also calculate and output the patient's eye information (including eye movement angle and fixation point); and the microprocessing system combines the information of the patient's three-dimensional model and eye information to output the degree of exophthalmos of the patient, as well as disease recognition based on facial state information (including expression, color, etc.). It can be understood that the above operations are all run on a computing PC processor, and the patient information collected at the same time is encrypted to prevent privacy leakage. Combining the above output information and calculation algorithms, facial multi-modal information such as RGBD images, videos, and audio is output, providing rich diagnostic data for doctors. Further, the data synchronization structure topology diagram is as Figure 4 shown, realized through PTP network synchronization, and can achieve a synchronization accuracy of up to hundreds of microseconds, ensuring high synchronization of multi-modal data collection.

[0059] In this embodiment, as Figure 1 shown, the facial multi-modal information acquisition device further includes: a fixed platform 120 and a lifting bracket 130; the fixed platform 120 is fixedly connected to the hemispherical shell 110 in the unfolded state, and is used to carry the hemispherical shell 110 and the RGBD camera system, the eye tracking system, the lighting system, and the microprocessing system installed on the hemispherical shell 110. The lifting bracket 130 is fixedly connected to the fixed platform 120, and is used to support the fixed platform 120 and drive the fixed platform 120 to reciprocate in the vertical direction, so as to adjust the height of the fixed platform 120 and the hemispherical shell 110, and further enable the opening 1101 of the hemispherical shell 110 to face the faces of patients with different heights, so as to be able to collect the facial modal information of patients with different heights, and further enable the facial multi-modal information acquisition device to have better versatility.

[0060] On the other hand, this embodiment also provides a specific process for a patient to use the facial multi-modal information acquisition device, including:

[0061] 1. The patient performs identity authentication (swiping a card / swiping a QR code / face recognition) on the facial multi-modal information acquisition device, and the patient's identity information is synchronized to the microprocessing system. The microprocessing system enters the patient's basic information and medical record information;

[0062] 2. The patient independently undergoes the examination under the guidance of the facial multi-modal information acquisition device. The facial multi-modal information acquisition device asks the patient for information such as physical signs and medical history. The patient enters the information by voice and confirms or corrects the information through voice / touchpad. The confirmed patient education information is transmitted to the microprocessing system, and the microprocessing system judges the suspected disease types of the patient according to the CAS scoring standard.

[0063] 3. Before the examination, the facial multi-modal information acquisition device detects the current environmental brightness and color, and automatically compensates the light; adjusts the fine-tuning of the light intensity and color temperature according to the patient's age and eye conditions; locates the patient's head position, and guides the patient to correct the body position and head angle through voice.

[0064] 4. During the examination, the patient's head and shoulders enter from the left side, and under the voice guidance of the facial multi-modal information acquisition device, the patient swings and adjusts the head / eye position; the facial multi-modal information acquisition device records the patient's video in real time, and simultaneously acquires 16 pictures at the angles as shown in Figure 3 Figure. The requirements for these 16 pictures are as follows: ① Images of nine eye observation directions. When taking pictures of the lower right, lower, and lower left directions, the eyelids need to be pulled open so that the eyeballs can be better photographed; ② Images of the patient's face turned 45° and 90° to the left and right respectively; ③ Frontal closed-eye image and frontal downward-looking image of the patient's face; ④ Upright image of the patient's face.

[0065] Specifically, the image shooting logic follows the following rules: Shoot in three groups according to whether the eyeballs move and whether the eyes are closed; the first group is to obtain the upright, 45° and 90° left and right images at one time when the eyeballs do not move, a total of 5 pictures; the second group is the frontal closed-eye and frontal downward-looking images, a total of 2 pictures; the third group is the images of nine eye observation directions, a total of 9 pictures.

[0066] More specifically, the working principle of the camera to shoot images: To improve the image acquisition efficiency and reduce the light interference of multiple cameras taking pictures, a certain order is adopted within one working cycle; the specific order is as follows:

[0067] (1) The first round of image acquisition

[0068] ① At time sequence 1, the five RGBD cameras located on the second semi-circular arc sequentially acquire RGBD images. The sixth RGBD camera exposes B - the seventh RGBD camera exposes J, the sixth RGBD camera transmits B - the first RGBD camera A exposes, the seventh RGBD camera transmits J - the eighth RGBD camera I exposes, the first RGBD camera A transmits - the ninth RGBD camera C exposes, the eighth RGBD camera I transmits - the ninth RGBD camera C transmits; where exposure means taking pictures, and transmission means transmitting the images to the microprocessing system.

[0069] ② At timing 2, the five RGBD cameras located on the first semi-circular arc collect images in sequence. The second RGBD camera D exposes - the third RGBD camera F exposes, the second RGBD camera D transmits - the first RGBD camera A exposes, the third RGBD camera F transmits - the fourth RGBD camera G exposes, the first RGBD camera A transmits - the fifth RGBD camera E exposes, the fourth RGBD camera G transmits - the fifth RGBD camera E transmits;

[0070] ③ At timing 3, if the patient has an eye examination item, the eye movement camera H takes an eye image and transmits it to the microprocessing system;

[0071] (2) Multiple rounds of image collection are performed in the order of the first round of image collection until the number of qualified images in each orientation exceeds 5.

[0072] (3) Using the depth camera 3D point cloud technology, by emitting infrared rays with a wavelength of 780 nanometers, the axial length of the eye and other eye parameters are measured without contacting the eyes, and the data is synchronized to the microprocessing system.

[0073] 5. The RGBD camera system transmits the images to the microprocessing system. The microprocessing system performs quality control on the 16 groups of eye position images that have been collected. According to the standard image library and the image quality control algorithm, pictures with incorrect ranking, position deviation, motion artifacts, and non-medical foreign objects are automatically removed; if all the images in a certain group do not meet the requirements, the patient is prompted to re-take the images of that group through the RGBD camera system until the images of that group meet the requirements, and the microprocessing system finally retains the images after passing the quality control.

[0074] 6. The doctor reads and diagnoses the films through the ophthalmic examination equipment. During the diagnosis process, the patient information, education information, and eye measurement information can be retrieved, and the images and videos of the eyes, orbits, and faces are measured and analyzed for 3D reconstruction, and a diagnosis report is output.

[0075] In summary, the innovation of the present invention lies in proposing a brand-new facial multi-modal information collection device, which realizes comprehensive and precise detection of the patient's face and eyes by integrating a variety of advanced collection technologies. The present invention not only solves the timing matching problem in existing multi-modal information collection, improves the comprehensiveness and accuracy of orbital disease diagnosis, but also provides a powerful data basis for artificial intelligence automatic diagnosis, with important clinical application value and promotion prospects. In addition, based on the facial multi-modal information collection device, the present invention also realizes the simultaneous and co-location collection of multi-modal information through an innovative process, further solving the timing matching problem existing in the prior art, thereby improving the diagnostic accuracy of orbital diseases and the accuracy of treatment plan decision-making, and greatly improving the treatment effect and quality of life of patients.

[0076] Although the content of the present invention has been described in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present invention. After those skilled in the art have read the above content, various modifications and alternatives to the present invention will be obvious. Therefore, the protection scope of the present invention should be defined by the appended claims.

Claims

1. A facial multi-modal information acquisition device for the diagnosis and treatment of orbital diseases in patients, characterized in that, Comprising: A hemispherical housing; The opening of the hemispherical housing faces the patient's face, and the patient's head is located at the center of the sphere of the hemispherical housing; An RGBD camera system, fixed to the inner surface of the hemispherical housing, for capturing RGBD images of the patient's face and shooting videos; each pixel point in the RGBD image includes R, G, B color information and depth information; An eye tracking system, fixed to the inner surface of the hemispherical housing, for tracking and obtaining the rotation angle of the patient's human eyes and the gaze direction; An illumination system, fixed to the inner surface of the hemispherical housing, for providing the required brightness when capturing RGBD images and videos and guiding the patient to observe the eyeballs; An audio input system, fixed to the inner surface of the hemispherical housing, for collecting audio data of the doctor and the patient; And A microprocessing system, communicatively connected to the RGBD camera system, for measuring and analyzing the RGBD images and the videos to generate a three-dimensional model of the patient's face.

2. The facial multi-modal information acquisition device according to claim 1, characterized in that The hemispherical housing includes a first semi-circular arc and a second semi-circular arc, and the plane where the first semi-circular arc is located is a horizontal plane, and the plane where the second semi-circular arc is located is a vertical plane perpendicular to the horizontal plane; the centers of the first semi-circular arc and the second semi-circular arc are both the center of the sphere of the hemispherical housing; the first semi-circular arc is tangent to the second semi-circular arc, and the tangent point is the midpoint of both; The RGBD camera system includes: 9 RGBD cameras; the 9 RGBD cameras are respectively denoted as the first RGBD camera, the second RGBD camera, the third RGBD camera, the fourth RGBD camera, the fifth RGBD camera, the sixth RGBD camera, the seventh RGBD camera, the eighth RGBD camera, and the ninth RGBD camera; Among them, the first RGBD camera is located at the tangent point, the second to fifth RGBD cameras are sequentially arranged on the first semi-circular arc, and the sixth to ninth RGBD cameras are sequentially arranged on the second semi-circular arc; and the second RGBD camera and the fifth RGBD camera are respectively located at the two end points of the first semi-circular arc, the third RGBD camera is located between the second RGBD camera and the first RGBD camera, and the fourth RGBD camera is located between the first RGBD camera and the fifth RGBD camera; the sixth RGBD camera and the seventh RGBD camera are located between the first end point of the second semi-circular arc and the first RGBD camera, and the seventh RGBD camera is close to the first RGBD camera; the eighth RGBD camera and the ninth RGBD camera are located between the first RGBD camera and the second end point of the second semi-circular arc, and the eighth RGBD camera is close to the first RGBD camera.

3. The facial multi-modal information acquisition device according to claim 2, wherein The sixth RGBD camera, the center of the sphere of the hemispherical housing, and the ninth RGBD camera form a first angle, the vertex of the first angle is the center of the sphere of the hemispherical housing, and the degree of the first angle is 110 degrees to 130 degrees.

4. The facial multi-modal information acquisition device according to claim 2, wherein, The eye tracking system includes: An eye movement camera, fixed on the second semi-circular arc, between the first RGBD camera and the eighth RGBD camera; An eye tracking display, fixed on the second semi-circular arc, between the first RGBD camera and the eye movement camera.

5. The facial multi-modal information acquisition device according to claim 4, characterized in that, The first RGBD camera, the center of the hemispherical housing, and the eye movement camera form a second angle. The vertex of the second angle is the center of the hemispherical housing, and the degree of the second angle is 20 degrees to 40 degrees.

6. The facial multi-modal information acquisition device according to claim 2, wherein The lighting system includes: A first lighting source and a second lighting source, located above the plane of the first semi-circular arc, fixed on the inner side of the hemispherical housing, for providing the required brightness when shooting RGBD images and videos; and the first lighting source and the second lighting source are respectively located on both sides of the plane of the second semi-circular arc; Nine LED lights, evenly arranged in 3 rows and 3 columns on the inner side of the hemispherical housing, for guiding the patient to observe the orientation of the eyeball.

7. The facial multi-modal information acquisition device according to claim 4, wherein, The audio input system includes at least one microphone; the microphone is fixed on the inner side of the hemispherical housing and is at the same horizontal height as the eye tracking display.

8. The facial multi-modal information acquisition device according to claim 1, wherein A headrest for supporting the patient's head is provided at the center of the hemispherical housing.

9. The facial multi-modal information acquisition device according to claim 1, wherein, The hemispherical housing has a contracted state and an expanded state; when diagnosing and treating orbital diseases of the patient, the hemispherical housing is in the expanded state; when not diagnosing and treating orbital diseases of the patient, the hemispherical housing is in the contracted state; and the hemispherical housing is made of light-tight material.

10. The facial multi-modal information acquisition device according to claim 9, wherein, It further includes: A fixed platform, fixedly connected to the hemispherical housing in the expanded state, for carrying the hemispherical housing; A lifting bracket, fixedly connected to the fixed platform, for supporting the fixed platform and driving the fixed platform to reciprocate in the vertical direction.