Traditional Chinese medicine remote diagnosis method and system based on low-delay transmission
By locating the anchor points of the nasal root and zygomatic process in the TCM remote diagnosis system, constructing a closed gradient phase-locked loop, and dynamically adjusting the encoding compression, the problem of information loss in key areas of TCM diagnosis caused by network fluctuations was solved, and the integrity and real-time performance of diagnostic information under low-latency transmission were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI BAOFANG TECHNOLOGY CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-21
AI Technical Summary
Existing TCM remote diagnosis systems suffer from ineffective encoding and compression strategies that fail to protect complexion and color information in key facial areas during network fluctuations, thus affecting the integrity of diagnostic information.
By locating the relevant anchor points of the nasal root and zygomatic process, the excitation base point is determined, a closed gradient phase-locked loop and color gradient abrupt change envelope region are constructed, the deformation tuning coefficient is calculated, and the video encoding compression degree is dynamically adjusted to ensure that the complexion features of key areas are not distorted during low-latency transmission.
It achieves the protection of complexion details in key areas of traditional Chinese medicine diagnosis under low-latency transmission conditions, avoids excessive encoding compression, and ensures the integrity of diagnostic information and real-time interaction.
Smart Images

Figure CN122437840A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real-time communication technology, and in particular to a method and system for remote diagnosis in traditional Chinese medicine based on low-latency transmission. Background Technology
[0002] In traditional Chinese medicine (TCM) remote diagnosis systems, existing technologies often employ general low-latency transmission protocols (such as WebRTC and SRT) combined with adaptive coding strategies. These strategies dynamically adjust video coding parameters based on network bandwidth, packet loss rate, or transmit buffer length to control end-to-end latency. These methods are primarily designed for general video call scenarios, and their coding parameter adjustments are mostly based on overall motion amplitude, global texture complexity, or buffer occupancy, with less consideration given to the specific color gradation boundary fidelity requirements of TCM diagnosis for particular anatomical areas (such as the root of the nose and the area near the cheekbone). In environments with frequent network fluctuations, the encoder may employ compression strategies to maintain low latency, such as increasing quantization parameters, reducing the frame rate, or selectively dropping frames. While this approach generally maintains low transmission latency, it may result in some loss of color and tone information in key facial areas.
[0003] Specifically, when network conditions change, the encoding and compression adjustment process rarely differentiates based on the real-time geometric features of the regions of interest in visual diagnosis (especially the color gradient abrupt envelope regions). In practical applications, the system may apply uniform compression parameters to the entire frame instead of prioritizing the protection of local boundaries and color gradient continuity that are of reference value for TCM diagnosis. Under certain conditions, this may manifest as a decrease in the clarity of certain facial color transition areas in the video stream received by the remote TCM doctor, thus affecting the integrity of the diagnostic information. Although this can be mitigated by increasing bandwidth or reducing resolution, these two methods each have different trade-offs under low latency constraints.
[0004] For example, a middle-aged female patient, suffering from chronic fatigue accompanied by faint reddish spots with indistinct borders on her cheekbones (preliminary diagnosis by local Traditional Chinese Medicine practitioners leaned towards "Yin deficiency and internal heat"), initiated a remote consultation with the hospital's geriatrics department via her home broadband (actual available uplink bandwidth fluctuating between 600kbps and 1.3Mbps). Approximately 25 seconds into the session, due to system updates on other devices in her home, the uplink bandwidth briefly dropped to approximately 450kbps. At this point, to maintain an end-to-end latency of no more than 250 milliseconds, the existing low-latency transmission system automatically changed the quantization parameters of the video encoding from... The latency was increased from 30 to 44, and some non-reference frames were discarded. At the receiving end of the TCM doctor at the remote end, noticeable color block merging and slight artifacts appeared in the transition area between the light red spots on the patient's cheeks and the surrounding normal skin color. Details that were originally helpful in judging the "degree of blurring of the edge of the spot" became difficult to distinguish. The TCM doctor could not determine whether the blurring phenomenon was due to the patient's true physical signs or transmission distortion. Therefore, the real-time observation was suspended, and the patient was asked to re-acquire and send static images. This extra operation increased the overall diagnostic latency from about 220 milliseconds to nearly 1.9 seconds, which to some extent offset the expected effect of low-latency transmission. Summary of the Invention
[0005] This invention provides a method and system for remote diagnosis in traditional Chinese medicine based on low-latency transmission, which avoids excessive compression of encoding caused by network fluctuations while maintaining low-latency transmission.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a remote diagnostic method for traditional Chinese medicine based on low-latency transmission, the method comprising: Step 1: Respond to the visual diagnosis session request and collect a continuous video frame sequence from the patient's camera as the raw bitstream to locate three anchor points: the center of the nasal root depression, the inflection points of the left and right zygomatic process edges, and so on. Based on the three anchor points, determine the first and second dual-domain pivot centers and determine whether the three anchor points are in the same straight line orientation to determine the activation base point. Step 2: Using the determined excitation base point as the origin, emit a set of search rays at equal angular intervals in the video frame plane. Each ray moves along the corresponding direction until it encounters a pixel position where the first directional derivative of the pixel gray value is flipped. Connect the corresponding pixels on all rays in sequence to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color level abrupt change envelope region. Step 3: Within the color gradient abrupt change envelope region, starting from the excitation base point, connect the two intersection points of each adjacent two rays with the two points of the closed gradient phase-locked loop trace with straight line segments to obtain a set of radial chord segments; construct a polygonal closed loop based on the set of radial chord segments; calculate the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt change envelope region to obtain the deformation tuning coefficient. Step 4: Combine the deformation tuning coefficient with the inter-frame variation of the video to dynamically adjust the compression level of the video encoding. Then, transmit the adjusted encoded video stream and the facial complexion feature vector extracted from the video stream back to the patient and the TCM doctor via a low-latency transmission protocol.
[0007] Secondly, a remote TCM diagnostic system based on low-latency transmission includes: The triggering base point discrimination module is used to respond to the visual diagnosis session request, collect the continuous video frame sequence of the patient's camera as the raw bit stream, and locate three stationary point markers: the center of the nasal root depression, the inflection points of the left and right zygomatic process edges; based on the three stationary point markers, determine the first and second dual-domain pivot centers, and determine whether the three stationary point markers are in the same straight line orientation to determine the triggering base point. The color gradient mutation envelope construction module is used to emit a set of search rays at equal angular intervals in the video frame plane with a determined excitation base point as the origin. Each ray moves along the corresponding direction until it encounters a pixel position where the first directional derivative of the pixel gray value is flipped. The pixels at the corresponding positions on all rays are connected in sequence to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color gradient mutation envelope region. The deformation tuning coefficient calculation module is used to connect the two intersection points of each pair of adjacent rays with the two points of the closed gradient phase-locked loop trace within the color gradient abrupt change envelope region, starting from the excitation base point, with straight line segments to obtain a set of radial chord segments; construct a polygonal closed loop based on a set of radial chord segments; calculate the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt change envelope region to obtain the deformation tuning coefficient. The low-latency bidirectional transmission module combines the deformation tuning coefficient with the inter-frame variation of the video to dynamically adjust the compression level of the video encoding. It then transmits the adjusted encoded video stream and the facial complexion feature vector extracted from the video stream back to the patient and the TCM doctor via a low-latency transmission protocol.
[0008] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0009] The above-described solution of the present invention has at least the following beneficial effects: This invention achieves precise localization of the core areas of facial visual diagnosis by locating anchor points related to the nasal root and zygomatic process and determining excitation base points. This avoids the loss of complexion and contour details in key areas of visual diagnosis caused by global compression, ensuring the integrity of the core features required for TCM diagnosis. By constructing closed gradient phase-locked loops and color gradient abrupt change envelope regions, it automatically identifies key areas of facial complexion abrupt changes, providing accurate regional basis for differentiated coding and ensuring that details such as complexion transitions and boundary contours with diagnostic value in TCM visual diagnosis are given priority protection. By calculating deformation tuning coefficients, it combines the morphological features of sensitive areas of facial visual diagnosis with changes between video frames to dynamically adjust the degree of encoding compression. While maintaining low-latency transmission, it avoids over-compression of encoding due to network fluctuations, reducing distortion of complexion details and transmission artifacts. By synchronously transmitting the adjusted encoded video stream and facial complexion feature vectors through a low-latency transmission protocol, it ensures real-time interaction between doctors and patients, effectively avoiding diagnostic interruptions caused by the inability to distinguish between image distortion and real physical signs. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating a remote diagnostic method for traditional Chinese medicine based on low-latency transmission, provided by an embodiment of the present invention.
[0011] Figure 2 This is a schematic diagram of a remote TCM diagnostic system based on low-latency transmission, provided by an embodiment of the present invention.
[0012] Figure 3 This is a schematic diagram of the location of the stationary marker and the determination of the excitation base point provided in the embodiment of the present invention.
[0013] Figure 4 This is a schematic diagram illustrating the effect of constructing a closed gradient phase-locked loop trace according to an embodiment of the present invention.
[0014] Figure 5 This is a trend chart of adaptive adjustment of dynamic compression ratio and quantization step size provided by an embodiment of the present invention.
[0015] Figure 6 This is a statistical chart showing the accuracy of facial complexion RGB feature vector extraction provided in an embodiment of the present invention. Detailed Implementation
[0016] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0017] like Figure 1As shown, embodiments of the present invention propose a method for remote observation diagnosis in traditional Chinese medicine based on low-latency transmission, the method comprising the following steps: Step 1: Respond to the visual diagnosis session request and collect a continuous video frame sequence from the patient's camera as the raw bitstream to locate three anchor points: the center of the nasal root depression, the inflection points of the left and right zygomatic process edges, and so on. Based on the three anchor points, determine the first and second dual-domain pivot centers and determine whether the three anchor points are in the same straight line orientation to determine the activation base point. Step 2: Using the determined excitation base point as the origin, emit a set of search rays at equal angular intervals in the video frame plane. Each ray moves along the corresponding direction until it encounters a pixel position where the first directional derivative of the pixel gray value is flipped. Connect the corresponding pixels on all rays in sequence to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color level abrupt change envelope region. Step 3: Within the color gradient abrupt change envelope region, starting from the excitation base point, connect the two intersection points of each adjacent two rays with the two points of the closed gradient phase-locked loop trace with straight line segments to obtain a set of radial chord segments; construct a polygonal closed loop based on the set of radial chord segments; calculate the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt change envelope region to obtain the deformation tuning coefficient. Step 4: Combine the deformation tuning coefficient with the inter-frame variation of the video to dynamically adjust the compression level of the video encoding. Then, transmit the adjusted encoded video stream and the facial complexion feature vector extracted from the video stream back to the patient and the TCM doctor via a low-latency transmission protocol.
[0018] In this embodiment of the invention, the invention achieves precise localization of the core areas of facial observation by locating the anchorage points related to the nasal root and zygomatic process and determining the excitation base point. This avoids the loss of complexion and contour details in key areas of facial observation caused by global compression, ensuring the integrity of the core features required for TCM diagnosis. By constructing a closed gradient phase-locked loop and a color gradient abrupt change envelope region, the invention automatically identifies key areas of facial complexion abrupt changes, providing precise regional basis for differentiated coding and ensuring that details such as complexion transitions and boundary contours with diagnostic value in TCM observation are given priority protection. By calculating the deformation tuning coefficient, the morphological features of sensitive areas of facial observation are combined with changes between video frames to dynamically adjust the degree of encoding compression. While maintaining low-latency transmission, this avoids over-compression of encoding due to network fluctuations, reducing distortion of complexion details and transmission artifacts. The adjusted encoded video stream and facial complexion feature vector are synchronously transmitted through a low-latency transmission protocol, ensuring real-time interaction between doctors and patients and effectively avoiding diagnostic interruptions caused by the inability to distinguish between image distortion and real physical signs.
[0019] In a preferred embodiment of the present invention, step 1 involves responding to a visual diagnosis session request and acquiring a continuous video frame sequence from the patient's camera as the raw bitstream to locate three anchor points: the center of the nasal root depression, the inflection points of the left and right zygomatic edges, and so on. The geometric midpoint of the line connecting the first and second anchor points is used as the first dual-domain pivot; the geometric midpoint of the line connecting the second and third anchor points is used as the second dual-domain pivot. The planar coordinate values of the first, second, and third anchor points are extracted to construct a third-order determinant, and the value of the third-order determinant is calculated. When the value of the third-order determinant is zero, it is determined that the three are in the same straight-line orientation, and the first dual-domain pivot is used as the activation base point. When the value of the third-order determinant is not zero, it is determined that the three are not in the same straight-line orientation, and the second dual-domain pivot is used as the activation base point. Specifically, this includes: Upon receiving a TCM remote diagnosis session request initiated by a patient through a designated interactive entry point, the system immediately initiates the request processing flow. First, it performs a comprehensive verification of the session request, checking core elements such as patient identity information, session initiation permissions, request timeliness, and transmission parameters to confirm the request is legal, valid, and meets the requirements for a remote diagnosis session. After successful verification, the system immediately performs in-depth analysis of the request, extracting key parameters such as session initiation time, patient basic information, diagnosis needs (e.g., focus on complexion, areas of focus), and transmission bandwidth requirements. Simultaneously, it quickly responds to the request, sending a notification to the patient indicating that the request has been received and video capture is initiating, ensuring the patient understands the session's progress. The system then synchronously initiates the patient's video capture process, continuously capturing a sequence of frontal video frames of the patient's face at a stable frame rate of 25 frames per second using a pre-compatible camera on the patient's end, forming a coherent raw video stream. In the acquired continuous video frames, the system focuses on identifying the core facial areas used in traditional Chinese medicine's diagnostic theory of facial complexion. Employing a facial feature recognition algorithm combined with pixel-level fine-tuning, it performs pixel-by-pixel retrieval and calibration of key facial anatomical structures within the video frames. This identifies three anatomically significant locations: the center of the nasal root depression, the inflection point of the left zygomatic process, and the inflection point of the right zygomatic process. These three locations are then marked as corresponding stationary point markers, each with clear pixel coordinate information. The stationary point marker corresponding to the center of the nasal root depression is uniformly defined as the first stationary point marker, the one corresponding to the inflection point of the left zygomatic process as the second stationary point marker, and the one corresponding to the inflection point of the right zygomatic process as the third stationary point marker. After clarifying the correspondence and coordinate values of the three stationary point markers, the coordinates of the first and second dual-domain pivots are calculated based on the pixel coordinates of each stationary point marker.
[0020] The first dual-domain pivot is the geometric midpoint of the line connecting the first and second stationary point markers, i.e., the horizontal coordinate of the first dual-domain pivot = (horizontal coordinate of the first stationary point marker + horizontal coordinate of the second stationary point marker) / 2; the vertical coordinate of the first dual-domain pivot = (vertical coordinate of the first stationary point marker + vertical coordinate of the second stationary point marker) / 2; the second dual-domain pivot is the geometric midpoint of the line connecting the second and third stationary point markers, i.e., the horizontal coordinate of the second dual-domain pivot = (horizontal coordinate of the second stationary point marker + horizontal coordinate of the third stationary point marker) / 2; the vertical coordinate of the second dual-domain pivot = (vertical coordinate of the second stationary point marker + vertical coordinate of the third stationary point marker) / 2; using the above coordinate calculation formulas, the precise calculation of the two dual-domain pivots is completed, obtaining the reference pivot positions with stable positions and clear geometric meanings.
[0021] After calculating the coordinates of the two dual-domain pivots, the system extracts the horizontal and vertical coordinates of the three stationary point markers in the video frame plane, substitutes the three sets of coordinates into the matrix, and constructs a third-order determinant for determining the collinearity of the three points: ; The determinant is calculated using the diagonal expansion rule. The specific formula is: Determinant value = (Horizontal coordinate of the first stationary point × Vertical coordinate of the second stationary point × 1) + (Vertical coordinate of the first stationary point × 1) × (Horizontal coordinate of the third stationary point + 1) × (Horizontal coordinate of the second stationary point × Vertical coordinate of the third stationary point) - 1 × (Vertical coordinate of the second stationary point × Horizontal coordinate of the third stationary point) - (Vertical coordinate of the first stationary point × Horizontal coordinate of the second stationary point × 1) - (Horizontal coordinate of the first stationary point × 1) × (Vertical coordinate of the third stationary point). The system determines the spatial arrangement of the three stationary points based on this determinant value. If the calculated determinant value is zero, the three stationary points are considered to be in the same straight line, and the first dual-domain pivot is designated as the activation base point. If the calculated determinant value is not zero, the three stationary points are considered to be in different straight lines, and the second dual-domain pivot is designated as the activation base point. Adaptively selecting the activation base point based on the actual facial point arrangement allows subsequent region search and feature extraction to better match the patient's actual facial structure.
[0022] This embodiment identifies three anchor points—the center of the nasal root depression and the inflection points of the left and right zygomatic processes—to pinpoint the core facial regions crucial for diagnosis in Traditional Chinese Medicine (TCM) visual inspection. Using a geometric midpoint method to determine the dual-domain pivot allows for rapid acquisition of reference points, ensuring efficient subsequent image processing. By combining third-order determinant values to determine if the three anchor points are collinear, the excitation base points can be adaptively selected based on the patient's actual facial morphology, ensuring the base point positions more closely match the true facial structure. This method abandons the global positioning approach commonly used in video processing, providing reliable reference support for subsequent color gradation region analysis and transmission encoding adjustments. Simultaneously, it avoids the loss of key diagnostic features due to improper base point selection, ensuring the integrity of facial key region information during remote inspection.
[0023] In a preferred embodiment of the present invention, step 2 involves emitting a set of search rays at equal angular intervals within the video frame plane, using the determined excitation base point as the origin. Each ray moves along its corresponding direction until it encounters a pixel position where the first directional derivative of the pixel's grayscale value undergoes a sign flip. The pixels at corresponding positions on all rays are then sequentially connected to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color-gradient abrupt change envelope region, including: Step 200a: Using the pixel coordinates of the excitation base point in the video frame plane as the origin of the polar coordinate system, the angular step size is set to π / 180 radians. Starting from 0 radians, the angular step size increases incrementally until 2π radians, generating the direction angles of each ray sequentially. Specifically, after accurately determining the excitation base point, the system establishes a polar coordinate system adapted for facial diagnostic key area search, centered on the pixel coordinates corresponding to the excitation base point in the video frame plane, and fixes the excitation base point as the origin of this polar coordinate system; To address the computational efficiency limitations in low-latency transmission scenarios, the system uniformly sets the increment step size of the ray direction angle to π / 180 radians. Starting from a 0-radian position, the system incrementally increases the direction angle according to this step size, continuously generating search rays with different directions until the direction angle completely covers the entire 2π-radian range. This completes the generation of search rays with full-angle and uniform distribution, ensuring a full-area scan of the core facial diagnostic area without omissions, overlaps, or blind spots, and avoiding missed detections of boundary points due to uneven ray distribution.
[0024] Step 201a: For each ray, starting from the excitation base point, the system advances one pixel at a time along the corresponding ray direction, calculating the first-order directional derivative of the current pixel's grayscale value relative to the ray direction. Simultaneously, the system calculates the same directional derivative of the previous pixel. Specifically, for each search ray with a defined direction angle, the system starts from the origin of the excitation base point and advances the search pixel by pixel along the ray's direction, with each step limited to one pixel unit to ensure the precision and accuracy of facial color boundary localization. During continuous stepping, the system simultaneously calculates the grayscale change rate (i.e., the first-order directional derivative) between the current and previous pixels along the ray direction. First, the color pixels in the video frame are grayscaled, extracting the red, green, and blue channel brightness values for each pixel. The grayscale value is calculated using the standard grayscale weighted formula: Grayscale value = 0.299 × red channel value + 0.587 × green channel value + 0.114 × blue channel value, where 0.299, 0.587, and 0.114 are weighting coefficients. The first-order directional derivative of a pixel along the ray direction is calculated as the ratio of the grayscale difference between adjacent pixels to the step distance, i.e., first-order directional derivative = (current pixel grayscale value - previous pixel grayscale value) ÷ pixel step distance. Since the pixel step distance is fixed at a single pixel and has a value of 1, the first-order directional derivative can be directly simplified to the difference between the current pixel grayscale value and the previous pixel grayscale value, effectively reducing calculation time and system resource consumption.
[0025] Step 202a: Multiply the first-order directional derivative of the current pixel by the first-order directional derivative of the previous pixel. If the product is less than zero, it is determined that the first-order directional derivative at the current pixel has undergone a sign flip, the coordinates of the corresponding pixel are recorded, and the search for the corresponding ray is terminated. If the product is not less than zero, continue to the next pixel until all pixels along the corresponding ray direction have been traversed. Specifically, after sequentially obtaining the first-order directional derivative values of the current pixel and the previous pixel along the same ray direction, the system multiplies the two derivative values. The sign of the product is used to determine whether the grayscale change trend has undergone a sudden flip. If the product result is less than zero, it is determined that the first-order directional derivative at the current pixel has undergone a sign flip. If the two first-order directional derivative values are positive and negative with opposite signs, it indicates a significant grayscale change at the current pixel location. This location corresponds to the boundary or contour edge of facial complexion, which is a key feature that needs to be preserved in traditional Chinese medicine's observation diagnosis. The system immediately records the precise coordinates of the boundary pixel and stops the subsequent search of the current ray. If the product result is greater than or equal to zero, it indicates that the current location is still in a region of uniform grayscale change and has not reached the target boundary location of the complexion change. The system continues to step along the original ray direction to the next pixel, repeating the derivative calculation, numerical multiplication, and sign judgment process until all valid pixels along the ray direction have been traversed.
[0026] Step 200b involves sequentially connecting the corresponding pixels on all rays to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color gradient abrupt change envelope region. Specifically, after all search rays have completed boundary pixel searches and determined their corresponding grayscale abrupt change points, the system sequentially and orderly connects all recorded boundary pixels according to the ascending order of ray direction angles, forming a continuous, closed, non-intersecting, and unbroken contour trajectory. This trajectory is the closed gradient phase-locked loop. The internal connected region enclosed by the closed gradient phase-locked loop is formally defined as the facial color gradient abrupt change envelope region, which concentrates the pixel group with the most significant changes in facial complexion.
[0027] This embodiment constructs a uniformly distributed search ray with the excitation base point as the origin, which can locate the boundaries of abrupt changes in facial color grayscale, avoiding missegmentation of sensitive areas in visual diagnosis by general image processing methods. Boundary points are determined by flipping the sign of the first-order directional derivative; the computational logic is simple and efficient, enabling region localization in a very short time, meeting the real-time requirements of low-latency transmission in remote visual diagnosis. The resulting closed gradient phase-locked loop and color-level abrupt change envelope can pinpoint core areas for TCM diagnosis, such as the cheekbone and perinasal region, avoiding distortion of visual diagnosis information caused by network fluctuations and compression.
[0028] In a preferred embodiment of the present invention, step 3 involves, within the color gradient abrupt change envelope region, starting from the excitation base point, connecting the two intersection points of each pair of adjacent rays with the two points of the closed gradient phase-locked loop trace using straight line segments to obtain a set of radial chord segments; constructing a polygonal closed loop based on the set of radial chord segments; calculating the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt change envelope region to obtain the deformation tuning coefficients, including: Step 300a: Collect the intersection points of each search ray and the closed gradient phase-locked loop. Each ray has exactly one intersection point. Arrange all intersection points sequentially according to the ascending direction angle of each ray to form an ordered set of intersection points. Specifically, the system first performs a unified traversal and sorting of all completed rays to extract the boundary pixels of the facial color level change region. Specifically, the system extracts the pixel coordinates of the intersection of each ray and the closed gradient phase-locked loop, and these coordinates are the boundary nodes of the color level change. Since each ray extends along a fixed single angle direction and only stops when a gray level change is detected, there will only be one valid intersection point between each ray and the closed gradient phase-locked loop, and there will be no abnormal situations of multiple intersection points or no intersection points. The system follows the direction angle sorting rules when generating rays and sorts all extracted intersection points in ascending order of direction angle to form an ordered set of intersection points with continuous positions and progressive angles. During the sorting process, the excitation base point is always the center, and the principle of increasing angle is followed to ensure that the arrangement of intersection points is consistent with the distribution of rays.
[0029] Step 301a: Take the first and second intersection points from the ordered set of intersection points, and connect the two intersection points with a straight line segment using the excitation base point as the viewpoint to generate the first radial chord segment; then take the second and third intersection points in sequence and connect them with a straight line segment to generate the second radial chord segment; repeat this process until the penultimate intersection point and the last intersection point are connected to obtain a set of radial chord segments. Specifically, the system extracts the pixel coordinates of two adjacent sets of intersection points in sequence from the ordered set of intersection points that have been sorted in ascending order of azimuth angles, and determines the completed excitation base point as the observation center that remains unchanged throughout the process. The adjacent intersection points are directly connected with a straight line segment to generate the first radial chord segment. After completing the construction of the first radial chord segment, the system continues to follow the predetermined arrangement order of the ordered set of intersection points, reads the subsequent adjacent intersection points one by one and connects them with straight line segments, generates the subsequent radial chord segments one by one, and continues to perform the adjacent intersection point connection operation until the second to last intersection point in the set is connected with the last intersection point with a straight line. Finally, a complete radial chord segment is formed with the excitation base point as the center and uniformly radiating along the fixed ray direction generated in step 200a. This set of radial chord segments can truly restore the internal contour undulation state and boundary extension trend of the determined color gradient abrupt envelope region.
[0030] Step 300b: Arrange a set of radial chord segments in the order they were generated, so that the end of the first radial chord segment coincides with the beginning of the second radial chord segment, the end of the second radial chord segment coincides with the beginning of the third radial chord segment, and so on, so that the end of each radial chord segment becomes the beginning of the next radial chord segment, forming a non-closed broken-line chain. Specifically, the system arranges and connects all radial chord segments in the order they were actually generated, and during the connection process, the coordinates of the end of the first radial chord segment coincides with the coordinates of the beginning of the second radial chord segment. The coordinates of the end of the second radial chord segment coincide with the coordinates of the beginning of the third radial chord segment. This process is repeated to connect all radial chord segments sequentially, so that the end of each radial chord segment directly serves as the beginning of the next radial chord segment. This results in a continuous and complete broken line chain with tightly connected nodes, extending sequentially along a fixed angle direction that increases by π / 180 arc degrees, but without achieving closure. This broken line chain fully preserves the concave and convex changes and morphological features of the original boundary of the color gradient abrupt envelope region in step 200b, intuitively reflecting the true structure of the key areas in traditional Chinese medicine observation diagnosis.
[0031] Step 301b: Using the excitation base point as the ring center and a radial span of three pixels as the radial distance, construct an equidistant pivot reference ring. For each linear segment in the non-closed polyline chain, calculate the intersection point between the corresponding linear segment's line and the equidistant pivot reference ring. Retain the intersection point located on the line connecting the midpoint of the corresponding linear segment and the ring center, and whose distance from the midpoint is the shortest, as the tangential calibration node of the corresponding linear segment. Specifically, the system uses the completed excitation base point as the fixed ring center and sets the radial span to a fixed 3 pixels, which serves as the constant radius of the equidistant pivot reference ring, thus constructing an equidistant pivot reference ring with a fixed center position, a uniform radius of 3 pixels, and a uniform and regular circumference. For each independent linear segment in the non-closed polyline chain formed in step 300b, first establish the linear equation based on the pixel coordinates of the two endpoints of the segment, and then solve the equation of the circle of the equidistant pivot reference ring simultaneously to calculate the intersection point of the linear segment and the reference ring. Specifically, taking the excitation base point as the origin, let the general equation of the linear segment be Ax + By + C = 0, and the circle equation of the equidistant pivot reference ring be x² + y² = 3. 2 =9. Solve the system of equations for the line and the circle simultaneously to obtain one or two sets of real solutions, corresponding to the coordinates of the intersection point of the line and the circle. Among all intersection points, select and retain target points that simultaneously meet two conditions: first, they must be located along a fixed line connecting the midpoint of the linear segment and the center of the circle; second, they must have the shortest pixel distance to the midpoint of the segment. These intersection points are then designated as the tangential calibration nodes for the corresponding linear segment. This equidistant pivot reference ring calibration method effectively corrects minor local distortions in the broken line chain caused by pixel positioning errors, image noise, or slight vibrations of the acquisition equipment, eliminating the impact of local morphological deviations on the accuracy of subsequent geometric calculations and parameter extraction.
[0032] Step 302b involves connecting all tangential calibration nodes sequentially according to their order in the non-closed polyline chain, forming a transition polyline passing through each tangential calibration node. The last tangential calibration node is then connected to the first tangential calibration node using a linear segment, closing the transition polyline and forming a closed polygonal loop with its ends connected. Specifically, the system connects all tangential calibration nodes calculated in step 301b strictly according to their original order in the non-closed polyline chain, along a fixed angle arrangement direction that increases uniformly by π / 180 radians, ensuring a smooth transition between adjacent calibration nodes and forming a transition polyline without breaks, abrupt changes, or intersections / twisting. After completing the overall transition polyline construction, the system further connects the last tangential calibration node to the first tangential calibration node in the transition polyline using a straight line segment, forming a complete closed structure and ultimately creating a closed polygonal loop with its ends connected, nodes evenly distributed, and a regular contour. This polygonal closed loop eliminates image noise, pixel positioning deviations, and local minor distortions present in the original boundary of the closed gradient phase-locked loop trace in step 200b through tangential calibration, and can objectively, stably, and accurately characterize the overall geometric shape of the color gradient abrupt change envelope region.
[0033] Step 300c: Obtain the boundary curve of the polygonal closed loop, perform arc length integration on the boundary curve to obtain the perimeter of the polygonal closed loop; obtain the color gradient abrupt change envelope region enclosed by the closed gradient phase-locked loop trace, perform area integration on the color gradient abrupt change envelope region to obtain the area of the color gradient abrupt change envelope region. Specifically, this includes: extracting the complete boundary curve of the polygonal closed loop, performing arc length integration on the boundary curve using a pixel-by-pixel numerical integration method, assuming that the polygonal closed loop is formed by connecting nodes P1(x1,y1), P2(x2,y2), ..., Pn(xn,yn) in sequence, and the Euclidean distance between two adjacent points is the differential segment length Li, and accumulating all differential segment lengths along the curve extension direction to finally calculate the overall perimeter of the polygonal closed loop. At the same time, the system extracts the complete color gradient abrupt change envelope region enclosed by the closed gradient phase-locked loop trace in step 200b, and also performs area integration on this region using a pixel-by-pixel numerical integration method. Traversing all valid pixels within the region in row and column order, with the area of a single pixel unit as the micro-area, and the total number of valid pixels in the region being N, and the area of a single pixel denoted as S0, the overall area of the color gradient abruptness envelope region is calculated using the formula: S = N × S0. This is then accumulated sequentially to obtain the overall area of the color gradient abruptness envelope region. Both integral calculations are based on the actual pixel size of the video frame to ensure that the obtained perimeter and area values closely match the true geometric scale of the image.
[0034] Step 301c: Divide the perimeter of the polygonal closed loop by the area of the color-gradient abrupt change envelope region to obtain a basic ratio; weight and fuse the basic ratio with the eddy current perturbation factor to obtain the deformation tuning coefficient. The process of obtaining the eddy current perturbation factor is as follows: extract a set of radial string segments, take the length of each radial string segment as the amplitude value of the eddy current signal, and take the difference in the plane angle between two adjacent radial string segments at the excitation base point as the phase offset; calculate an eddy current perturbation factor characterizing the degree of local morphological distortion within the color-gradient abrupt change envelope region based on the amplitude values of all radial string segments and the adjacent phase offsets, specifically including: Dividing the perimeter of the polygonal closed loop by the area of the color gradient abrupt change envelope yields a basic ratio that reflects the compactness, contour complexity, and morphological regularity of the region. This basic ratio directly reflects the geometric characteristics of the region: with similar areas, a larger perimeter indicates a more tortuous contour boundary with more variations in concavity and convexity, corresponding to higher contour complexity, lower morphological regularity, and poorer regional compactness; conversely, with similar perimeters, a smaller area indicates a more compact and concentrated region with relatively lower contour complexity and higher morphological regularity. Therefore, the value of the basic ratio K0 directly quantifies the compactness, contour tortuosity, and overall morphological regularity of the facial color gradient abrupt change envelope.
[0035] After obtaining the basic ratio, this basic ratio is weighted and fused with the eddy current disturbance factor to finally obtain the deformation tuning coefficient that can comprehensively characterize the dynamic changes in the regional morphology, i.e., K = α × K0 + β × The weighting coefficient α is 0.6 and β is 0.4, satisfying α+β=1. This not only highlights the overall geometric features of the region represented by the base ratio, but also takes into account the local distortion details reflected by the eddy current disturbance factor, achieving a reasonable balance between the two types of features. The eddy current perturbation factor is used to quantitatively characterize the degree of local distortion, unevenness, and irregularity within the color gradient abrupt change envelope region. The larger the value, the more dramatic the undulation of the regional boundary and the more obvious the local distortion. The smaller the value, the smoother the region's outline and the more regular its shape. The specific method for obtaining the eddy current disturbance factor is as follows: the system extracts all radial chord segments generated in step 301a, and calculates the pixel length of each radial chord segment in the video frame. As the amplitude value of the eddy current signal, the difference in the plane angle formed by two adjacent radial chord segments at the excitation base point is used. As a phase offset, the amplitude value and the phase offset are first normalized separately, i.e. , In the formula It is the first The dimensionless amplitude value of the normalized length of the spoke-to-chord segment. It is the maximum value among all radial chord segment pixel lengths. It is the dimensionless phase offset after normalization. It is the maximum value among all adjacent angle differences. Then, variance statistics and fusion calculations are performed on the normalized values. ,in, It is the total number of radial chord segments. This is the arithmetic mean of all product terms. This formula directly characterizes the degree of local distortion, unevenness, and irregularity within the color-gradient abrupt change envelope region by calculating the fluctuation variance of the amplitude and phase shift of adjacent radial chord segments. A larger value indicates a more chaotic distribution of chord lengths and angles, corresponding to severe unevenness and higher degrees of local distortion in the region; a smaller value indicates a more uniform and smoother contour, and better local morphological regularity. The resulting deformation tuning coefficients simultaneously integrate the overall geometric compactness of the region with the details of local distortion fluctuations. They can be directly used as core control parameters to guide the dynamic adjustment of subsequent video coding compression intensity, prioritizing the integrity of the complexion features and image clarity of key diagnostic areas under the premise of low-latency transmission.
[0036] This embodiment constructs radial chord segments and broken line chains through ordered intersections, objectively restoring the true shape of facial color gradient change areas, avoiding irrelevant interference, and matching the characteristic description requirements of key areas such as the cheekbones and perinasal region in traditional Chinese medicine diagnosis. Tangential calibration using equidistant pivot reference rings eliminates contour distortion caused by image noise and pixel positioning errors, making polygonal closed loops more stable and reliable. A basic ratio constructed using the perimeter-to-area ratio can quantify the geometric compactness of key diagnostic areas; the calculation method is simple and efficient, without affecting the real-time requirements of low-latency transmission. Eddy current disturbance factors effectively reflect local concavity and convexity distortion, allowing deformation tuning coefficients to simultaneously consider overall shape and local details, better reflecting the actual characteristics of facial color changes. Deformation tuning coefficients provide a direct and targeted adjustment basis for video coding compression, prioritizing the protection of color details in core diagnostic areas during network fluctuations, reducing transmission distortion.
[0037] In a preferred embodiment of the present invention, step 4 includes: Step 400: Extract the pixel luminance matrices of two adjacent frames in the current video frame sequence, calculate the Frobenius norm of the difference between the two matrices, and calculate the normalized inter-frame difference equivalent by dividing the Frobenius norm by the total number of pixels in a single frame. Specifically, this includes: in a continuous video frame sequence transmitted in real-time, first locating the target frame to be encoded, and simultaneously retrieving the previous frame (the reference frame that has already been encoded). Perform pixel-level luminance value analysis on these two frames to ensure that the luminance data of each pixel is accurately collected without missing any pixel's luminance information. The image is processed according to a fixed row and column resolution (set to 1920×1080 pixels, i.e., the total number of rows in the image). =1080 rows, total number of columns =1920 columns), starting from the top left pixel, the brightness value of each pixel in the two frames of the image after grayscale transformation is read row by row and column by column. According to the actual spatial position of the pixel in the image, the brightness matrix of the previous frame and the brightness matrix of the current target frame are constructed respectively. The number of rows and columns of the two matrices corresponds exactly to the image resolution (1080 rows × 1920 columns), and each element in the matrix corresponds to the brightness value of a pixel in the image, thus completely preserving the brightness distribution details of the two frames of the image.
[0038] To objectively and accurately measure the overall difference between two adjacent frames, the system first performs a difference operation on the corresponding elements of the previous frame's brightness matrix A and the current frame's brightness matrix B, calculating the brightness difference of each pixel point by point, thereby constructing a complete difference matrix D. Let the previous frame's brightness matrix be A and the current target frame's brightness matrix be B. The construction of the difference matrix D satisfies D = A - B, where the element in the I-th row and J-th column of the difference matrix D... The element corresponding to the brightness matrix A of the previous frame Elements of the current frame luminance matrix B The difference, i.e. = This element directly reflects the magnitude of the brightness change at the corresponding pixel location. A positive difference indicates that the pixel's brightness is higher in the current frame than in the previous frame, a negative difference indicates that the pixel's brightness is lower in the current frame than in the previous frame, and a difference of 0 indicates that the pixel's brightness has no change. The overall difference of the difference matrix D is calculated using the Frobenius norm. = Where 1080 is the total number of rows in the image and 1920 is the total number of columns. This refers to the Frobenius norm. A larger norm value indicates a more significant difference in brightness between two frames, signifying a greater range of image change; a smaller norm value indicates that the two frames are more similar, with a smaller range of change. The calculated Frobenius norm is divided by the total number of pixels in a single frame (1920×1080) to achieve dimensionless normalization of the difference value, eliminating numerical bias caused by different resolutions. This results in a normalized inter-frame difference equivalent. This normalized inter-frame difference equivalent value is calculated by standardizing the total brightness difference across the entire image with the total number of pixels. A larger value indicates a more significant change in average brightness per pixel, visually reflecting a larger overall movement and more obvious content switching between adjacent frames; a smaller value indicates a weaker change in brightness per pixel, visually reflecting a more static image with essentially unchanged content.
[0039] Step 401: Multiply the deformation tuning coefficients by the normalized inter-frame difference equivalent to obtain the dynamic compression ratio; using the preset base quantization step size as the initial value, perform a weighted fusion of the dynamic compression ratio and the base quantization step size to obtain the quantization step size offset; weight the base quantization step size and the quantization step size offset to obtain the corrected quantization step size for the current frame. Specifically, this includes: multiplying the obtained deformation tuning coefficients by the normalized inter-frame difference equivalent obtained in step 400 to obtain a dynamic compression ratio that can simultaneously adapt to the morphological features of facial color gradient change regions and the degree of inter-frame image change. Using the system-preset base quantization step size (value 16, unit: quantization unit) as a fixed initial reference, perform a weighted product operation on the dynamic compression ratio and the base quantization step size to obtain the quantization step size offset used for dynamically adjusting the encoding compression intensity; then perform a weighted sum operation on the original base quantization step size and the quantization step size offset, and determine the final calculation result as the corrected quantization step size actually used in the current video frame.
[0040] Step 402: Encode and compress the current video frame according to the corrected quantization step size to generate an adjusted encoded video stream; decode the reconstructed video frame from the adjusted encoded video stream, extract the pixel group located in the color gradient abruption envelope region of the reconstructed video frame, and calculate the mean vector of the pixel group in the red, green and blue channels as the facial complexion feature vector. Specifically, this includes: calling a dedicated encoding algorithm, using the corrected quantization step size as the core parameter, formally starting the encoding process, preprocessing the image data of the current video frame, reading the brightness and color values of each pixel in the image one by one, and organizing them in an orderly manner by row and column according to the original image size to construct a standardized data matrix. Perform Discrete Cosine Transform (DCT) to complete the conversion of image data from the spatial domain to the frequency domain. The preprocessed image data will be divided into processing units of 8×8 pixels (i.e., one image block), and each image block corresponds to a set of continuous pixel brightness data. For each 8×8 image block, a discrete cosine transform (DCT) operation is performed. The frequency domain coefficients are calculated as: DCT = 8×8 matrix × pixel brightness matrix × DCT transpose. The DCT matrix is a fixed 8×8 matrix, and the pixel brightness matrix is the brightness value matrix of the current image block. Through matrix multiplication, the pixel brightness data in the spatial domain is converted into coefficient data in the frequency domain, including low-frequency and high-frequency coefficients. Low-frequency coefficients range from 0 to 100, primarily corresponding to the overall image outline, such as the overall brightness and general shape of the image. All coefficients within this range are preserved to ensure that the overall image features required for traditional Chinese medicine diagnosis are not lost. High-frequency coefficients range from 101 to 255, corresponding to the details of the image, including the texture of areas with abrupt color changes, subtle brightness variations, and subtle differences in facial complexion. This effectively reduces data redundancy and transmission pressure while ensuring that key details are not lost, achieving a balance between compression efficiency and information integrity.
[0041] After completing the conversion from the spatial domain to the frequency domain, the frequency domain coefficients are subjected to hierarchical quantization processing according to the corrected quantization step size (based on the basic quantization step size of 16, combined with the specific value adjusted by the dynamic offset). For low-frequency coefficients with values from 0 to 100, all are retained to ensure the overall image outline is clear. For high-frequency coefficients with values from 101 to 255, the selection is flexible according to the amplitude of image changes. More are retained when the image changes greatly and the details are rich, and appropriate compression is applied when the image is flat. When the image is flat (normalized inter-frame difference equivalent ≤ 0.05), the part of the high-frequency coefficients with values from 101 to 150 is retained (corresponding to the subtle textures of facial complexion and the edge details of color gradation change areas, which are the core details of TCM observation), and the part with values from 151 to 255 is discarded (corresponding to irrelevant minor noise and non-critical textures). This effectively controls data redundancy, reduces transmission pressure, and accurately retains the key details required for TCM observation, avoiding the loss of diagnostic information such as complexion details and color gradation changes due to excessive compression. After quantization, the quantized frequency domain coefficients are further compressed and encoded using Huffman coding (a type of entropy coding) to remove redundant data and generate an encoded video stream adapted to the current video frame. The compression strength of this video stream is matched with the corrected quantization step size. When the image changes significantly, the compression strength is moderate, and when the image changes slightly, the compression strength is slightly higher, which satisfies the transmission efficiency requirements without losing core image information.
[0042] The generated encoded video stream is reverse-decoded to reconstruct complete video frames. The core features (including color gradation distribution, brightness changes, and detail textures) of the reconstructed frames are compared one by one with those of the original video frames to confirm that the color gradation change areas and key details related to facial complexion are not distorted or omitted, ensuring that the core information required for TCM diagnosis is fully preserved. After confirmation, the color gradation change envelope area in the reconstructed video frame is located, and the effective pixel group in this area is extracted, eliminating interference from irrelevant pixels. Then, the pixel values of this pixel group in the red, green, and blue channels are counted respectively. The average brightness value of each channel is calculated by averaging. The average brightness values of the three channels are then combined in a fixed order of red, green, and blue to form a complete facial complexion feature vector. The average brightness value of the red channel corresponds to the warmth or coolness of the facial complexion, the average brightness value of the green channel corresponds to the softness of the complexion, and the average brightness value of the blue channel corresponds to the clarity of the complexion. The combination of the three clearly reflects the core features of facial complexion and directly reflects the overall state of facial complexion. The encoded video stream and the color feature vector are uniformly encapsulated, integrated into a composite data unit according to a fixed format, and a data check code is added to ensure data integrity during transmission. Then, it is sent through a dedicated transmission channel to ensure real-time transmission.
[0043] Step 403: Package the adjusted encoded video stream and facial complexion feature vector into a composite data unit; send the composite data unit to the patient and TCM doctor via a low-latency transmission protocol through a hybrid transmission channel using User Datagram Protocol (UDP) and Real-Time Transmission Control Protocol (RTC). Specifically, this includes: organizing the data of the encoded video stream after adaptive encoding compression adjustment, clarifying the core data of the encoded video stream (including video frame data, encoding parameters, data verification information, etc.), and summarizing the extracted and calculated facial complexion feature vector (containing the average brightness values of the red, green, and blue channels, arranged in a fixed order), and performing unified data encapsulation processing on these two types of core data. The encapsulation process follows a pre-defined fixed data format. First, the encoded video stream data is standardized to ensure a uniform and non-redundant data format. Then, the facial complexion feature vector is associated and matched with the encoded video stream data, supplementing key information such as data checksums, data identification information, and transmission priority identifiers. Finally, it is integrated to form a well-structured, complete, and highly efficient composite data unit. This unit clearly divides the encoded video stream data area, the complexion feature vector area, and the data check area. Each area has clear identifiers and data format requirements to ensure that there are no errors or losses during data transmission, while significantly reducing the amount of data transmitted and adapting to remote transmission needs.
[0044] To fully meet the core requirements of low latency and high real-time performance in TCM remote diagnosis, the system adopts a hybrid transmission mode of User Datagram Protocol (UDP) and Real-time Transmission Control Protocol (RTCP) to build a stable and reliable transmission channel. UDP is primarily responsible for the rapid transmission of encoded video streams and core data such as complexion feature vectors, minimizing transmission latency and ensuring fast data delivery. RTCP is responsible for real-time monitoring of the transmission status, providing feedback on data transmission progress and integrity, and promptly identifying and correcting packet loss and errors during transmission to prevent data transmission interruptions. Simultaneously, relying on a dedicated low-latency transmission protocol, the system dynamically allocates transmission bandwidth, prioritizing the transmission speed of key TCM diagnostic information (complexion features, image details) to further reduce transmission latency and ensure a stable and efficient transmission process. Through this hybrid transmission channel, the encapsulated composite data units are synchronously sent to both the patient's display terminal and the TCM doctor's treatment terminal. The composite data units sent to the patient's display terminal primarily retain the video image and basic complexion information; the composite data units sent to the TCM doctor's treatment terminal fully retain all encoded video streams, complexion feature vectors, and data verification information, ensuring that the TCM doctor can clearly obtain core information such as facial complexion details and image changes. This achieves synchronous transmission and efficient interaction of encoded video streams and complexion feature data. The entire transmission process is monitored throughout, with real-time feedback on the transmission status, ensuring data transmission without delay or loss.
[0045] This embodiment constructs a dynamic compression ratio by combining inter-frame difference equivalents and deformation tuning coefficients. It can adaptively adjust the coding intensity according to the morphology and image change intensity of key facial diagnostic areas, balancing transmission efficiency and image quality. Encoding with a modified quantization step size can preserve the complexion details within the color gradient abrupt change envelope area to the greatest extent while ensuring low-latency transmission, avoiding distortion of key diagnostic features. Extracting complexion feature vectors directly from the reconstructed frames can complete the core feature extraction in advance, reducing the subsequent parsing pressure on the TCM doctor. Using a hybrid transmission of composite data units using User Datagram Protocol (UDP) and Real-Time Transmission Control Protocol (RTC) ensures both low latency in the video stream and reliable data transmission.
[0046] like Figure 2 As shown, embodiments of the present invention also provide a remote diagnostic system for traditional Chinese medicine based on low-latency transmission, comprising: The triggering base point discrimination module is used to respond to the visual diagnosis session request, collect the continuous video frame sequence of the patient's camera as the raw bit stream, and locate three stationary point markers: the center of the nasal root depression, the inflection points of the left and right zygomatic process edges; based on the three stationary point markers, determine the first and second dual-domain pivot centers, and determine whether the three stationary point markers are in the same straight line orientation to determine the triggering base point. The color gradient mutation envelope construction module is used to emit a set of search rays at equal angular intervals in the video frame plane with a determined excitation base point as the origin. Each ray moves along the corresponding direction until it encounters a pixel position where the first directional derivative of the pixel gray value is flipped. The pixels at the corresponding positions on all rays are connected in sequence to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color gradient mutation envelope region. The deformation tuning coefficient calculation module is used to connect the two intersection points of each pair of adjacent rays with the two points of the closed gradient phase-locked loop trace within the color gradient abrupt change envelope region, starting from the excitation base point, with straight line segments to obtain a set of radial chord segments; construct a polygonal closed loop based on a set of radial chord segments; calculate the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt change envelope region to obtain the deformation tuning coefficient. The low-latency bidirectional transmission module combines the deformation tuning coefficient with the inter-frame variation of the video to dynamically adjust the compression level of the video encoding. It then transmits the adjusted encoded video stream and the facial complexion feature vector extracted from the video stream back to the patient and the TCM doctor via a low-latency transmission protocol.
[0047] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0048] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0049] Experimental example: I. Experiment Overview Experimental conditions This experiment was conducted on a remote diagnostic system platform of the Traditional Chinese Medicine (TCM) department of a tertiary hospital. The experiment took place from March to May 2024, recruiting 30 healthy adult volunteers (aged 22-55, half male and half female) and 20 patients with typical facial complexion characteristics identified by TCM diagnosis (including sallow complexion, flushed cheeks, pale complexion, etc.). The experimental equipment included a 1080P high-definition camera for the patient, a 4K display terminal for the TCM practitioner, and a remote diagnostic server deployed in the hospital's data center.
[0050] The network environment simulated a real-world home broadband scenario: a base bandwidth of 1.0 Mbps, with periodic fluctuations introduced by the network simulator, the bandwidth varying between 600 kbps and 1.3 Mbps, a fluctuation period of 15-30 seconds, and an occasional packet loss rate of 0.5%-2%. Video encoding used the H.264 standard, with a base quantization step size of 30, a frame rate of 25 fps, and a resolution of 1920×1080. Comparison schemes included: traditional WebRTC adaptive encoding (fixed QP adjustment strategy), SRT transmission protocol, and the method described in this invention.
[0051] Step 1: Locating the station marker and determining the trigger point: Upon responding to the visual diagnosis session request, the system collects a continuous sequence of video frames from the patient's camera as the raw bitstream, with a frame rate set at 25fps and a resolution of 1920×1080 pixels. A facial feature recognition algorithm is used to locate three anchor points: the center of the nasal root depression, and the inflection points of the left and right zygomatic edges. In the experiment, anchor point localization tests were conducted on 50 subjects, and the average localization accuracy of the three anchor point markers reached the pixel level (error < 2 pixels).
[0052] Based on three stationary point markers, the first dual-domain pivot (the midpoint of the line connecting the first and second stationary points) and the second dual-domain pivot (the midpoint of the line connecting the second and third stationary points) are calculated. A third-order determinant is used to determine if the three stationary points are collinear: if the determinant is zero, they are considered collinear, and the first dual-domain pivot is used as the excitation base point; otherwise, the second dual-domain pivot is used. Experimental data shows that in normal facial structures, the three stationary points are not collinear in approximately 94% of cases, indicating that the adaptive excitation base point selection mechanism effectively covers the vast majority of practical application scenarios.
[0053] Figure 3The simulation results demonstrate the localization of stationary point markers and the determination of the excitation base point. The figure clearly marks the locations of three stationary points: the center of the nasal root depression (P1), the inflection point of the left zygomatic process edge (P2), and the inflection point of the right zygomatic process edge (P3), as well as the first dual-domain pivot, the second dual-domain pivot, and the finally determined excitation base point. The pixel coordinates of each stationary point, the calculation results of the dual-domain pivot, and the determinant determination value are all labeled in the figure, verifying the accuracy of the localization algorithm.
[0054] Step 2: Constructing a closed gradient phase-locked loop: Using a defined excitation point as the origin, search rays are emitted at equal angular intervals within the video frame plane. The angular step size is set to π / 180 radians (2 degrees), starting from 0 radians and extending to 2π radians, generating a total of 360 search rays covering the entire circumference. Each ray moves pixel-by-pixel along its corresponding direction, calculating the first-order directional derivative of the current pixel's grayscale value with respect to the ray direction. When the product of the first-order directional derivatives of adjacent pixels is less than zero, a sign flip is determined, and the pixel coordinates are recorded as a boundary point.
[0055] By connecting the boundary pixels recorded on all rays in ascending order of orientation angle, a closed gradient phase-locked loop is constructed. The region inside this loop is defined as the color gradient abrupt change envelope, which concentrates the pixel group with the most significant changes in facial complexion. In the experiment, facial images of 50 subjects were processed, and the success rate of constructing the closed gradient phase-locked loop was 100%, with an average processing time of 12.3 ms, meeting the real-time requirements of low-latency transmission.
[0056] Figure 4 The diagram illustrates the construction effect of a closed gradient phase-locked loop. Centered on the excitation point, 360 search rays are evenly distributed, each terminating at a boundary point where the grayscale gradient sign is flipped. Connecting all boundary points forms a closed gradient phase-locked loop (solid white line). The region inside the loop is the color-gradient abrupt change envelope. The diagram also shows grayscale change curves along some ray paths, clearly demonstrating the sign flipping of the first-order directional derivative at the boundary points.
[0057] Step 4: Dynamic Video Encoding and Low-Latency Transmission: By combining the deformation tuning coefficients with the inter-frame variation in video, the compression level of the video encoding is dynamically adjusted. The pixel brightness matrices of two adjacent frames are extracted, the Frobenius norm of the difference matrix is calculated, and the ratio of this to the total number of pixels in a single frame yields the normalized inter-frame difference equivalent. Multiplying the deformation tuning coefficients by the inter-frame difference equivalent gives the dynamic compression ratio, which is then weighted and fused with the base quantization step size to obtain the corrected quantization step size. The base quantization step size is set to 30, and the weighting coefficients are adaptively adjusted based on network conditions.
[0058] The current video frame is encoded and compressed according to the corrected quantization step size. Simultaneously, the RGB mean vector of pixel groups within the color gradient abrupt change envelope region is extracted from the reconstructed video frame as a facial complexion feature vector. The encoded video stream and the complexion feature vector are packaged into a composite data unit and sent to both the doctor and patient ends via a UDP+RTP hybrid transmission channel. Experimental data shows that, under network bandwidth fluctuations of 600kbps to 1.3Mbps, the system's average end-to-end latency is controlled within 185ms, a 26% reduction compared to traditional methods.
[0059] Figure 5 The diagram illustrates the adaptive adjustment trend of the dynamic compression ratio and quantization step size. The figure shows the processing of 100 consecutive frames of video. When the inter-frame difference increases (due to rapid motion), the dynamic compression ratio increases accordingly, and the quantization step size adaptively increases to maintain low latency. Conversely, when the inter-frame difference decreases, the dynamic compression ratio decreases, and the quantization step size decreases to preserve color details. The quantization step size is dynamically adjusted within the range of 28 to 42, with an average value of 34.2, ensuring both low-latency transmission and avoiding color distortion caused by over-compression.
[0060] Facial complexion feature extraction accuracy: Facial images of 50 subjects were used to extract complexion features. Using RGB values manually annotated by professional TCM practitioners as a benchmark, the accuracy of the facial complexion feature vectors extracted by the method of this invention was calculated. Experimental results showed that the average accuracy of the R channel was 96.8%, the average accuracy of the G channel was 95.4%, the average accuracy of the B channel was 94.2%, and the overall accuracy was 95.5%. Compared with the traditional global compression method (overall accuracy 87.3%), this represents an improvement of 8.2 percentage points.
[0061] Figure 6 The statistical results of facial complexion RGB feature vector extraction accuracy are presented. The figure compares the feature extraction accuracy of the proposed method with traditional global compression and fixed QP coding methods in the R, G, and B channels. The proposed method achieves the highest accuracy in all three channels, with 96.8% for the R channel, 95.4% for the G channel, and 94.2% for the B channel, significantly outperforming the comparison methods. This indicates that the proposed method effectively preserves the detailed information about facial complexion required for traditional Chinese medicine diagnosis through precise localization and differential coding of the color gradient abrupt change envelope region.
[0062] Through the detailed description and experimental verification of the above embodiments, the traditional Chinese medicine remote diagnosis method based on low-latency transmission described in this invention exhibits the following beneficial effects: By locating three anchor points—the center of the nasal root depression and the inflection points of the left and right zygomatic edges—and adaptively determining the excitation base point, precise localization of the core areas of facial observation was achieved. This avoided the loss of complexion and contour details in key areas of facial observation caused by global compression, ensuring the integrity of the core features required for TCM diagnosis. By constructing closed gradient phase-locked loops and color gradient abrupt change envelopes, key areas of facial complexion abrupt changes were automatically identified, providing accurate regional basis for differential coding. By calculating deformation tuning coefficients, the morphological features of sensitive areas of facial observation were combined with changes between video frames to dynamically adjust the degree of coding compression. While maintaining low-latency transmission (average 185ms), over-compression due to network fluctuations was avoided, reducing distortion of complexion details and transmission artifacts.
[0063] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for remote observation diagnosis in traditional Chinese medicine based on low-latency transmission, characterized in that, The method includes: Step 1: Respond to the visual diagnosis session request and collect a continuous video frame sequence from the patient's camera as the raw bitstream to locate three anchor points: the center of the nasal root depression, the inflection points of the left and right zygomatic process edges, and so on. Based on the three anchor points, determine the first and second dual-domain pivot centers and determine whether the three anchor points are in the same straight line orientation to determine the activation base point. Step 2: Using the determined excitation base point as the origin, emit a set of search rays at equal angular intervals in the video frame plane. Each ray moves along the corresponding direction until it encounters a pixel position where the first directional derivative of the pixel gray value is flipped. Connect the corresponding pixels on all rays in sequence to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color level abrupt change envelope region. Step 3: Within the color gradient abrupt change envelope region, starting from the excitation base point, connect the two intersection points of each adjacent two rays with the two points of the closed gradient phase-locked loop trace with straight line segments to obtain a set of radial chord segments; construct a polygonal closed loop based on the set of radial chord segments; calculate the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt change envelope region to obtain the deformation tuning coefficient. Step 4: Combine the deformation tuning coefficient with the inter-frame variation of the video to dynamically adjust the compression level of the video encoding. Then, transmit the adjusted encoded video stream and the facial complexion feature vector extracted from the video stream back to the patient and the TCM doctor via a low-latency transmission protocol.
2. The method for remote TCM diagnostic observation based on low-latency transmission according to claim 1, characterized in that, Based on three stagnation point markers, the first and second dual-domain pivot centers are determined. It is then determined whether the three stagnation point markers are aligned along the same straight line to identify the excitation base point, including: The geometric midpoint of the line connecting the first and second stationary points is taken as the first double-domain pivot; the geometric midpoint of the line connecting the second and third stationary points is taken as the second double-domain pivot; the planar coordinate values of the first, second, and third stationary points are extracted, a third-order determinant is constructed, and the value of the third-order determinant is calculated; When the value of the third-order determinant is equal to zero, it is determined that the three are in the same straight line orientation, and the first double-domain pivot is taken as the activation base point; when the value of the third-order determinant is not equal to zero, it is determined that the three are not in the same straight line orientation, and the second double-domain pivot is taken as the activation base point.
3. The method for remote TCM diagnostic observation based on low-latency transmission according to claim 2, characterized in that, Using a defined excitation point as the origin, a set of search rays are emitted at equal angular intervals within the video frame plane. Each ray moves along its corresponding direction until it encounters a pixel position where the sign of the first directional derivative of the pixel's grayscale value is flipped, including: Using the pixel coordinates of the excitation base point in the video frame plane as the origin of the polar coordinate system, the angular step size is set to π / 180 radians. Starting from 0 radians, the angular step size is increased until 2π radians, and the direction angles of each ray are generated sequentially. For each ray, starting from the excitation base point, step by one pixel distance along the corresponding ray direction each time, calculate the first directional derivative of the gray value of the current pixel with respect to the ray direction, and at the same time calculate the same directional derivative of the previous pixel. Multiply the first-order directional derivative of the current pixel with the first-order directional derivative of the previous pixel. If the product is less than zero, it is determined that the first-order directional derivative at the current pixel has been sign-flipped. Record the coordinates of the corresponding pixel and terminate the search of the corresponding ray. If the product is not less than zero, continue to step to the next pixel until all pixels in the corresponding ray direction have been traversed.
4. The method for remote TCM diagnostic observation based on low-latency transmission according to claim 3, characterized in that, Within the color gradient abrupt change envelope region, starting from the excitation base point, connecting the two intersection points of each pair of adjacent rays with the two points of the closed gradient phase-locked loop with straight line segments yields a set of radial chord segments, including: Collect the intersection points of each search ray and the closed gradient phase-locked loop. Each ray has one and only one intersection point. Arrange all the intersection points in order of increasing ray direction angle to form an ordered set of intersection points. Take the first and second intersection points from the ordered set of intersection points, and connect the two intersection points with a straight line segment using the excitation base point as the viewpoint to generate the first radial chord segment; then take the second and third intersection points in sequence and connect them with a straight line segment to generate the second radial chord segment; repeat this process until the penultimate intersection point and the last intersection point are connected to obtain a set of radial chord segments.
5. The method for remote TCM diagnostic observation based on low-latency transmission according to claim 4, characterized in that, Construct a polygonal closed loop based on a set of radial chord segments, including: Arrange a set of radial chord segments in the order they were generated, so that the end of the first radial chord segment overlaps with the beginning of the second radial chord segment, the end of the second radial chord segment overlaps with the beginning of the third radial chord segment, and so on, so that the end of each radial chord segment becomes the beginning of the next radial chord segment, forming a non-closed zigzag chain. With the excitation base point as the center and a length of three pixels as the radial span, an equidistant pivot reference ring is constructed. For each linear segment in the non-closed polyline chain, the intersection point between the corresponding linear segment and the equidistant pivot reference ring is calculated. The intersection point located on the line connecting the midpoint of the corresponding linear segment and the center of the ring, and which is the shortest distance from the midpoint, is retained as the tangential calibration node of the corresponding linear segment. Connect all tangential calibration nodes sequentially according to their order in the non-closed polyline chain to form a transition polyline passing through each tangential calibration node; take the last tangential calibration node and connect it with the first tangential calibration node with a linear segment to close the transition polyline and form a polygonal closed loop with the beginning and end connected.
6. The method for remote TCM diagnostic observation based on low-latency transmission according to claim 5, characterized in that, The ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt envelope region is calculated to obtain the deformation tuning coefficients, including: Obtain the boundary curve of the polygonal closed loop, perform arc length integration on the boundary curve to obtain the perimeter of the polygonal closed loop; obtain the color gradient abrupt change envelope region enclosed by the closed gradient phase-locked loop trace, perform area integration on the color gradient abrupt change envelope region to obtain the area of the color gradient abrupt change envelope region. Divide the perimeter of the polygonal closed loop by the area of the color gradient abrupt envelope region to obtain a basic ratio; then weight and fuse the basic ratio with the eddy current disturbance factor to obtain the deformation tuning coefficient.
7. The method for remote TCM diagnostic observation based on low-latency transmission according to claim 6, characterized in that, The process of obtaining the eddy current disturbance factor is as follows: Extract a set of radial string segments, take the length of each radial string segment as the amplitude value of the eddy current signal, and take the difference of the plane angle between two adjacent radial string segments at the excitation base point as the phase offset; calculate an eddy current perturbation factor that characterizes the degree of local morphological distortion in the color gradient abrupt envelope region based on the amplitude values of all radial string segments and the adjacent phase offsets.
8. The method for remote observation diagnosis in traditional Chinese medicine based on low-latency transmission according to claim 7, characterized in that, Step 4 includes: Extract the pixel brightness matrix of two adjacent frames in the current video frame sequence, calculate the Frobenius norm of the difference between the two matrices, and calculate the ratio of the Frobenius norm to the total number of pixels in a single frame to obtain the normalized inter-frame difference equivalent. The deformation tuning coefficients are multiplied by the normalized inter-frame difference equivalent to obtain the dynamic compression ratio; the dynamic compression ratio and the base quantization step size are weighted and fused to obtain the quantization step size offset; the base quantization step size and the quantization step size offset are weighted and used as the corrected quantization step size for the current frame. The current video frame is encoded and compressed according to the corrected quantization step size to generate an adjusted encoded video stream; the reconstructed video frame is decoded from the adjusted encoded video stream, and the pixel group located in the color gradient change envelope region in the reconstructed video frame is extracted. The mean vector of the pixel group in the red, green and blue channels is calculated as the facial complexion feature vector. The adjusted encoded video stream and facial complexion feature vector are packaged together to form a composite data unit. The composite data unit is then sent to the patient and TCM doctor via a low-latency transmission protocol through a hybrid transmission channel that combines User Datagram Protocol (UDP) with Real-Time Transmission Control Protocol (RTC).
9. A remote diagnostic system for traditional Chinese medicine based on low-latency transmission, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The triggering base point discrimination module is used to respond to the visual diagnosis session request, collect the continuous video frame sequence of the patient's camera as the raw bit stream, and locate three stationary point markers: the center of the nasal root depression, the inflection points of the left and right zygomatic process edges; based on the three stationary point markers, determine the first and second dual-domain pivot centers, and determine whether the three stationary point markers are in the same straight line orientation to determine the triggering base point. The color gradient mutation envelope construction module is used to emit a set of search rays at equal angular intervals in the video frame plane with a determined excitation base point as the origin. Each ray moves along the corresponding direction until it encounters a pixel position where the first directional derivative of the pixel gray value is flipped. The pixels at the corresponding positions on all rays are connected in sequence to construct a closed gradient phase-locked loop. The internal region of the closed gradient phase-locked loop is defined as the color gradient mutation envelope region. The deformation tuning coefficient calculation module is used to connect the two intersection points of each pair of adjacent rays with the two points of the closed gradient phase-locked loop trace within the color gradient abrupt change envelope region using straight line segments, starting from the excitation base point, to obtain a set of radial chord segments; based on a set of radial chord segments, a polygonal closed loop is constructed; The deformation tuning coefficient is obtained by calculating the ratio of the perimeter of the polygonal closed loop to the area of the color gradient abrupt envelope region. The low-latency bidirectional transmission module combines the deformation tuning coefficient with the inter-frame variation of the video to dynamically adjust the compression level of the video encoding. It then transmits the adjusted encoded video stream and the facial complexion feature vector extracted from the video stream back to the patient and the TCM doctor via a low-latency transmission protocol.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.