Skin cancer risk assessment method and system combining vision and language model
By cross-locating and sequentially arranging skin lesion images and medical record texts, the problem of inaccurate image-text fusion in existing technologies is solved, achieving high-precision identification for skin cancer risk assessment.
Patent Information
- Application Number
- CN202610168039.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-03-10
AI Technical Summary
In skin cancer risk assessment, existing technologies fail to provide detailed analysis of local color changes and edge closure trends in image processing, neglecting minor abnormalities in early lesions. Text processing fails to identify semantic repetitions and logical relationships between paragraphs, and image-text fusion lacks a location correspondence mechanism. This results in fragmented risk assessment criteria and reduces the specificity and coherence of lesion identification.
By collecting the color channel distribution and edge trends of skin lesion images, and combining them with symptom word nodes and time information in medical record text, cross-location and serialization of images and text are performed. A comparison mechanism between image structural jumps and text paragraphs is established to achieve accurate mapping and expression tendency arrangement of image and text information, and output skin cancer risk identification results.
It achieves precise labeling of lesion areas, improves the granularity of capturing changes in image structure, enhances the coherence of expression, ensures the consistency of information along the evolution path, clarifies the attribution of risk areas, and improves recognition accuracy.
Smart Images

Figure CN121639704A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and artificial intelligence, and in particular to a method and system for skin cancer risk assessment that combines visual and language models. Background Technology
[0002] The field of medical image processing and artificial intelligence technology involves analyzing medical images using computer vision and artificial intelligence algorithms. Core tasks include image acquisition, preprocessing, feature extraction, segmentation, recognition, and classification.
[0003] The traditional method and system for skin cancer risk assessment using a combined visual and language model involves integrating skin lesion images with medical record data. This is achieved through computer vision analysis of the images, natural language processing to extract textual information, and deep learning algorithms for data fusion and processing. This method enables multimodal data fusion, assisting clinicians in assessing skin cancer risk.
[0004] Image processing fails to analyze local color transitions and edge closure trends in detail, easily overlooking subtle abnormalities in early lesions. Text processing fails to identify semantic repetition and logical relationships between paragraphs, leading to information confusion. Image-text fusion lacks a location correspondence mechanism, making it impossible to establish an effective link between image structural changes and specific symptom expressions. In multimodal collaboration, there is no precise mapping between structural lines and temporal descriptions, making it difficult to establish a comparison between image and text content, resulting in fragmented risk assessment criteria and reduced specificity and coherence in lesion identification. Summary of the Invention
[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a method and system for skin cancer risk assessment using a combined visual and language model. The technical solution is as follows: A skin cancer risk assessment method combining visual and language models includes the following steps: S1: Collect the color channel distribution, edge trend direction and texture direction of skin lesion images, observe color changes and edge closure in the sliding window area, screen color concentration and structural continuity, number and process areas with closure features and color jumps, and output the visual change sequence of lesions. S2: Read the medical record text corresponding to the image, extract the symptom word nodes, time information and genetic tendency statements, split the paragraphs according to the position information, delete repeated expressions and rearrange the semantic order, determine whether there is a continuous expression relationship according to the lexical dependency structure, and output the medical record expression sequence content set. S3: Based on the color diffusion and structural boundary trend in the visual change sequence of the lesion, combined with the symptom words and time descriptions in the content of the medical record expression sequence, perform block cross-location on the image structural jump line and text segment, compare the image and text and classify them into partitions, and output the visual language joint expression content group. S4: Call the visual language joint expression content group, extract the connection mode of continuous segments of image color edges and text symptom words, compare whether the start and end ranges of the two expressions are continuous, process the combination items according to the image change path and language flow, and output the arrangement and integration sequence of expression tendency; S5: Call the image start position and text temporal symptom expression in the expression tendency arrangement integration sequence, track the color boundary extension area, compare the disease course paragraphs in the language expression, compare the continuous content nodes of the two types of sources, map the image block and language content to the risk area, and output the skin cancer risk identification result.
[0006] As a further aspect of the present invention, the lesion visual change sequence includes region number labels, color jump feature values, edge closure index, and structural continuity parameters; the medical record expression sequence content set includes a symptom node list, a time period annotation set, a genetic tendency segment group, and a semantic connection relationship table; the visual language joint expression content group includes image-text correspondence partitions, a structural jump control group, and symptom time intersection nodes; the expression tendency arrangement integrated sequence includes an image expression path sequence, a language structure sorting set, and expression coherence markers; and the skin cancer risk identification result includes a risk image block set, a disease course description paragraph group, and an image-text junction comparison table.
[0007] As a further aspect of the present invention, the step of obtaining the visual change sequence of the lesion is as follows: S101: Acquire images of skin lesions, read the pixel distribution of the red, green and blue color channels in the image, extract the color mutation locations based on the standard deviation and mean of the pixel values in the channel area, and obtain the coordinate set of the color mutation locations; S102: Based on the set of coordinates of the color change location, a sliding observation area is set in the image, and the degree of edge closure is determined according to the relationship between the gray gradient direction of the edge pixels in the area and the neighborhood distribution, and a numerical table of closed edge intervals is obtained. S103: Call the closed edge interval value table, combine the range of color channel value changes and texture direction change trends within the region, filter out regions that meet the conditions of structural closure and color jump, and number them to obtain the visual change sequence of the lesion.
[0008] As a further aspect of the present invention, the step of obtaining the medical record expression sequence content set is as follows: S201: Based on the medical record text data corresponding to the image, extract the symptom word nodes, time period information and genetic tendency sentence fragments, call the position index of various information in the text, split the sentence segments containing relevant content, and determine whether their position and semantic structure are repeated, and generate a set of word and time sentence segment decompositions. S202: Call the vocabulary and time segment decomposition set, sort the positions according to the segment index and time order, identify the vocabulary dependency structure between adjacent segments, and arrange the segments with related structures in sequence to obtain the dependency relationship segment arrangement sequence; S203: Based on the sequence of dependent relationship sentence segments, identify the repetition and synonymous structures of adjacent sentence segments, determine the degree of repetition by combining position offset information, perform deletion or merging operations on the repetitive sentence segments, and obtain the medical record expression sequence content set.
[0009] As a further aspect of the present invention, the step of obtaining the visual language joint expression content group is as follows: S301: Call the color diffusion part and structural boundary trend in the visual change sequence of the lesion, map the color value gradient change with the boundary line direction, compare the difference between the boundary angle change and the color diffusion edge, extract the points where the boundary trend and color diffusion trend are consistent, and generate a set of structural boundary fitting points. S302: Based on the start and end times of the time segment containing the symptom words in the content set of the medical record expression sequence, and using the visual frame number within the corresponding time period as a benchmark, extract the structural position within the corresponding frame in the set of fitted points, compare it with the time range of the text segment, and obtain the cross-temporal positioning segment value. S303: Based on the cross-temporal positioning segment value and the position trajectory of the structural jump line in the image, combined with the regional description information of symptom words in the text, the structural number and frame sequence number segments with consistent image and text positioning are selected and integrated to form a visual language joint expression content group.
[0010] As a further aspect of the present invention, the step of obtaining the expression tendency arrangement integrated sequence is as follows: S401: Based on the regional expression items in the visual language joint expression content group, extract the continuous pixel gradient region of the color edge in the image, split it into different line segments according to the spatial position, call the disease words in the text segment, determine the start and end positions of each word in the sentence, and obtain the image and text start and end association range by mapping the image edge segment to the text segment words through sequential mapping. S402: Based on the aforementioned image and text start and end association range, extract the path trend of the image edge line segment, calculate the positional order of continuous segments and turning segments in the line segment, call the word order structure of the text symptom words, determine whether the relationship between adjacent segments in the image path and the language flow is consistent, and obtain the corresponding expression order. S403: Based on the corresponding expression order, and combining the dense relationship of edge line segments in the image region with the distribution order of disease words, the index order is set according to the corresponding connection strength between the image and the text. The connection method between the regional expression items and the language structure is arranged in sequence to obtain the expression tendency arrangement integration sequence.
[0011] As a further aspect of the present invention, the steps for obtaining the skin cancer risk identification result are as follows: S501: Based on the image starting position in the integrated sequence of the expression tendency arrangement and the expression content of time symptoms in the text, extract the starting pixel point of the corresponding area of the image and the corresponding position of the time words in the paragraph, complete the position matching according to their correspondence in the image sequence and the text order, and obtain the index corresponding to the time image; S502: Call the image region and disease course expression paragraph associated in the time image corresponding index, identify the extension direction of the color boundary in the image, and compare it with the arrangement order of the disease course content in the language. Based on the relative position of the two, match the boundary region with the paragraph word order to obtain the image and text sequence correspondence result. S503: Based on the image boundary region and disease expression node in the image-text sequence correspondence result, extract the texture change path in the image and combine it with the concentrated location of the disease content to complete the joint regioning of the image and text, and filter the relevant region content according to the control range to obtain the skin cancer risk identification result.
[0012] A skin cancer risk assessment system combining visual and language models, the system comprising: The lesion visual acquisition module acquires images of skin lesions, reads the distribution of red, green and blue channels and marks the color concentration areas, detects edge contours and records the edge trend direction, reads the surface texture direction and forms directional labels, sets a sliding window and observes the color changes and edge closure status within the window in turn, screens closed feature areas and performs numbering processing on color jump areas, and generates a sequence of lesion visual changes. The medical record text reconstruction module obtains the medical record text data corresponding to the image, extracts symptom word nodes and marks their position in the sentence segment, extracts time period information and marks the start and end positions, extracts genetic tendency sentence fragments and marks family history keywords, disassembles text paragraphs and deletes duplicate paragraphs, rearranges the semantic order of paragraphs and marks continuous expression relationships, and generates a medical record expression sequence content set. The image-text cross-location module calls the color diffusion area and structural boundary trend in the visual change sequence of the lesion, calls the symptom vocabulary fragments and time description positions in the content set of the medical record expression sequence, marks the area where the image structure jump line is located and marks the position of the text description sequence segment, performs cross-location of image blocks and text fragments and assigns corresponding numbers, puts them into a unified partition and records the pairing relationship, and generates a visual language joint expression content group. The tendency arrangement integration module calls the image expression items and text expression items of each partition in the visual language joint expression content group, extracts continuous segments of color edges in the image and records their extension paths, extracts the connection mode of symptom words in the text segment and records the order of sentence flow, compares the connection relationship between the image extension path and the order of text flow and marks the consistency status, filters partitions with coherent consistency status and performs front and back sorting processing, merges the sorting results and forms the expression tendency arrangement index, and generates the expression tendency arrangement integration sequence. The risk identification and assessment module, based on the image start position and text temporal symptom expression content in the integrated sequence of expression tendency, tracks the color boundary extension area and marks the extension direction and coverage, tracks the disease course description paragraphs and marks their start and end time and symptom change description, performs image and text node mapping and merges the image blocks and language structure after the nodes to the risk control range, and generates skin cancer risk identification results.
[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this invention, color transitions and edge closure of images are dynamically read through sliding regions to achieve precise marking of lesion areas and improve the granularity of image structural changes capture. Medical record text is rearranged after removing redundant paragraphs, and symptom words and time information are extracted by combining semantic dependency relationships, effectively enhancing the coherence of expression. Block-level comparisons are established between image structural transitions and symptom fragments to achieve image-text position mapping. Sequential arrangement is performed based on the start and end ranges of image and text expression to ensure consistency of information along the evolutionary path. By extending boundaries and cross-connecting with disease progression segments, the attribution of risk areas is clarified, improving recognition accuracy. Attached Figure Description
[0014] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart illustrating the acquisition of the lesion visual change sequence according to the present invention. Figure 3 This is a flowchart illustrating the process of obtaining the medical record expression sequence content set according to the present invention. Figure 4 This is a flowchart illustrating the process of obtaining the visual language joint expression content group of the present invention; Figure 5 This is a flowchart illustrating the process of obtaining the integrated sequence of expression tendencies in this invention. Figure 6 This is a flowchart illustrating the process of obtaining skin cancer risk identification results according to the present invention. Detailed Implementation
[0015] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0016] refer to Figures 1 to 6 A skin cancer risk assessment method combining visual and language models includes the following steps: S1: Acquire images of skin lesions, read the color channel distribution, edge trend direction and surface texture direction of the images, use the sliding area to observe color changes and identify the degree of edge closure, screen the degree of color concentration and structural continuity trend, number the areas with closure characteristics and color jump performance, and output the visual change sequence of lesions. S2: Read the medical record text data corresponding to the image, extract the word nodes, time period information and genetic tendency sentence fragments related to symptoms, break down the content in the text paragraphs according to the positional relationship, delete duplicate paragraphs and rearrange the semantics, determine whether there is a continuous expression relationship through the word dependency structure, and output the medical record expression sequence content set. S3: Call the color diffusion part and structural boundary trend in the visual change sequence of the lesion, combine the symptom vocabulary in the content of the medical record expression sequence with the start and end positions of the time description, perform block cross-location on the structural jump line in the image and the descriptive sequence in the text, compare the image and text sources and classify them into the corresponding partition, and output the visual language joint expression content group. S4: Call the expression items of each region in the visual language joint expression content group, extract the connection between the continuous segments of color edges in the image and the words of symptoms in the text segment, compare whether the start and end ranges of the two are continuous in the expression process, perform pre- and post-processing on each combination item according to the image information transformation path and language structure flow, and output the expression tendency arrangement integrated sequence. S5: Invoke the image start position and the time-related symptom expression content in the integrated sequence of expression tendency, track the extended area of color boundary and paragraphs containing disease course content in language expression, perform junction comparison and mapping on the continuous content from the two sources, merge the image blocks and language structure after junction into the risk control range, and output the skin cancer risk identification result.
[0017] The visual change sequence of lesions includes region number labels, color jump feature values, edge closure index, and structural continuity parameters. The medical record expression sequence content set includes a symptom node list, time period annotation set, genetic tendency segment group, and semantic connection relationship table. The visual-language joint expression content group includes image-text correspondence partitions, structural jump control group, and symptom-time intersection nodes. The expression tendency arrangement integrated sequence includes image expression path sequence, language structure sorting set, and expression coherence marker. The skin cancer risk identification results include risk image block set, disease course description segment group, and image-text junction comparison table.
[0018] Please see Figure 2 The steps for obtaining the visual change sequence of the lesion are as follows: S101: Acquire images of skin lesions, read the pixel distribution of the red, green and blue color channels in the image, extract the color mutation locations based on the standard deviation and mean of the pixel values in the channel area, and obtain the coordinate set of the color mutation locations; For suspected melanoma skin lesion images with a resolution of 1024×1024 pixels, a depth color channel separation task was performed. The original image was split into three independent color matrices: red, green, and blue. The red channel data, which is most sensitive to the characteristics of subepidermal hemoglobin and melanin deposition, was explicitly used for core analysis. During the process, a 200×200 pixel region at the geometric center of the image was selected as the initial region of interest. By traversing the grayscale values of all points within this region pixel by pixel, the average grayscale value of this region was calculated to be 145.60, and the standard deviation reflecting the grayscale dispersion was 22.40. Based on the statistical regularity of the contrast between normal skin and lesion edges in clinical dermoscopic images, the system precisely set the baseline coefficient for color aberration judgment to 1.5, and calculated the upper threshold for judging color anomalies to be 179.20 and the lower threshold to be 112.00. When the system scans a pixel at coordinates (50, 60) with a grayscale value of 108.00, it compares the value with a lower threshold and determines that the pixel exceeds the grayscale fluctuation range of normal skin, thus marking it as a color mutation location coordinate. The system continuously performs precise comparison and filtering of such grayscale jump amplitudes on all pixels within the ROI area, eliminating background noise interference, and extracts the spatial coordinates of all discrete points that meet the mutation conditions, compiling them into a set of color mutation location coordinates that can characterize the initial contour of the lesion edge.
[0019] S102: Based on the coordinate set of color change location, a sliding observation area is set in the image. The degree of edge closure is determined according to the relationship between the gray-level gradient direction of the edge pixels in the area and the neighborhood distribution, and a numerical table of closed edge intervals is obtained. The system retrieves each coordinate point from the set of coordinates for color abrupt changes and sets a 5×5 pixel sliding observation area centered on that coordinate in the image space. During the sliding window coverage, the system uses the Sobel gradient operator to perform convolution operations on the central pixel within the window, obtaining its horizontal and vertical gradient components, and then calculates the gradient magnitude and 0.78 radians of that point. To determine the physical structural integrity of the lesion edge, the system further analyzes the distribution relationship of neighboring pixels within the 5×5 area, counting the number of edge pixels that satisfy the gradient continuity condition within an 8-neighborhood, resulting in an actual connected pixel count of 18. Simultaneously, the system performs geometric fitting based on the radius of curvature of the local area, calculating the theoretical total number of circumferential pixels required to form this closed edge segment as 20. By dividing the actual number of connected pixels by the theoretical total, the system obtains the closure index of this local edge as 0.90, and quantizes and verifies it against the preset structural closure judgment threshold of 0.85. The execution action continuously slides and iterates in the image. The system records key texture direction parameters in real time, such as the average gradient magnitude of 185.40, the closure index of 0.90, and 45.00 degrees. By summarizing and deduplicating the calculation results of each observation window, a numerical table of closed edge intervals reflecting the geometric closure of the lesion edge is obtained.
[0020] S103: Call the closed edge interval value table, combine the range of color channel value changes and texture direction change trends within the region, filter regions that meet the conditions of structural closure and color jump, and number them to obtain the visual change sequence of the lesion.
[0021] The system retrieves spatial and texture parameters from the closed edge interval value table and combines them with the red channel color distribution data extracted from S101 to verify the severity of color abrupt changes in the lesion area. During the screening process, the system calculates the difference between the maximum value (190.00) and the minimum value (108.00) of the pixels within the area, resulting in a color channel value variation range of 82.00. Simultaneously, it statistically analyzes the directional properties of the texture within the area, calculating a texture direction variance of 0.12. The above calculation results are then matched against a clinically established color abrupt change baseline of 30 and a texture consistency threshold of 0.50. If the color change range is greater than 30 and the texture direction variance is less than 0.50, the area is deemed to possess pathologically significant structural closure and color abrupt change characteristics. The system automatically numbers the eligible areas; for example, the first significantly abnormal area to pass the screening is marked as Seq_01, and the areas are sequentially arranged according to the spatial order from the top left to the bottom right of the image coordinate system. Through this process, the system accurately identifies visually aberrant units with high diagnostic value from complex background interference and obtains a sequence of visual changes in lesions that quantifies the irregular distribution characteristics of the lesions.
[0022] Please see Figure 3The steps for obtaining the medical record expression sequence content set are as follows: S201: Based on the medical record text data corresponding to the image, extract the symptom word nodes, time period information and genetic tendency sentence fragments, call the position index of various information in the text, split the sentence segments containing relevant content, and determine whether there is a repetition in their position and semantic structure, and generate a set of word and time sentence segment decompositions. Based on the electronic medical record text data corresponding to the images, a pre-trained named entity recognition model is used to extract symptom-related vocabulary nodes, time period information, and sentence fragments reflecting genetic risk, such as "black spots," "enlargement," and "family history." During execution, the system uses text position indexing technology to accurately record the start position (35) and end position (36) of the word "enlargement" in the original text stream, and logically splits long sentences into paragraphs based on these index numbers. To ensure the uniqueness of the information, the system performs semantic vectorization processing on each segment after splitting, and evaluates the semantic overlap by calculating the cosine similarity between vectors. If the similarity value reaches 0.95 or higher, the system will determine that the semantic structure is repeated and perform a merging operation. A second position check is performed on the sentence segments containing relevant content to ensure that each extracted medical event has a clear time anchor and position mark. Through refined parsing and cleaning of unstructured text, a set of vocabulary and time segment decompositions containing key pathological clues is generated.
[0023] S202: Call the vocabulary and time segment decomposition set, sort the positions according to the segment index and time order, identify the vocabulary dependency structure between adjacent segments, and arrange the segments with related structures in sequence to obtain the segment arrangement sequence of dependency relationship; The system retrieves a set of decomposed lexical and temporal sentence segments and arranges them in ascending order along the timeline based on specific dates or the sequence of disease progression descriptions within the time period information. After sorting, the system uses a dependency parsing algorithm to identify the lexical dependency structure between adjacent sentence segments, focusing on analyzing the modification and modified relationships between adjectives, verbs, and core nouns. For example, it identifies "range" as an attribute node for "spots," and "expansion" as an action node describing the dynamic evolution of "range." The system constructs logically oriented dependency chains, reorganizing discrete textual descriptions into phrase sequences that conform to medical logic, and then sequentially arranging sentence segments with related structures. This process eliminates random interference in the textual narrative, ensuring a high degree of closure in the symptom descriptions across both temporal and logical dimensions, thereby obtaining a sequence of dependency relationship sentence segments that clearly reflects the evolution of the patient's disease over time.
[0024] S203: Based on the sequence of dependent relationship sentence segments, identify the repetition and synonymous structures of adjacent sentence segments, determine the degree of repetition by combining position offset information, perform deletion or merging operations on repetitive sentence segments, and obtain the medical record expression sequence content set.
[0025] Based on the sequence of dependent relationship sentence segments, the system utilizes a medical ontology knowledge base to identify semantic repetitions and near-synonyms in adjacent sentence segments. For example, it identifies a high degree of semantic consistency between "increased diameter" and "expanded range" in clinical descriptions. During execution, the system calculates the semantic distance between different expressions and compares this distance with a preset near-synonym threshold of 0.20. It also incorporates character offset information within the text to comprehensively determine the degree of repetition. For identified repetitive expressions, the system performs deletion or merging operations based on the principle of maximizing information content, ensuring that each retained node contains unique diagnostic information. Through further refinement of the text's logical chain, the system transforms the chaotic medical record narrative into a standardized set of feature descriptions, resulting in a set of well-structured, logically coherent medical record expression sequences that have been freed of all redundant information.
[0026] Please see Figure 4 The steps for obtaining the visual language joint expression content group are as follows: S301: Call the color diffusion part and structural boundary trend in the visual change sequence of the lesion, map the color value gradient change with the boundary line direction, compare the difference between the boundary angle change and the color diffusion edge, extract the points where the boundary trend and color diffusion trend are consistent, and generate a set of structural boundary fitting points. The system retrieves the color diffusion component and geometric trend data of the structural boundary from the visual change sequence of the lesion, and performs spatial mapping and dot product operations on the color value gradient change vector of each pixel and the tangential vector of the boundary line. During execution, the system calculates the cosine value of the angle between the two vectors to quantify the consistency between the color diffusion direction and the boundary expansion trend, and sets a consistency comparison threshold of 0.96. When the system detects that the cosine value of the angle at a certain boundary point is 0.9986, since this value exceeds the threshold, the system determines that the boundary growth at that location is entirely driven by internal pigment diffusion and has extremely high clinical significance. By traversing all boundary segments in the image and performing this consistency check, the system extracts all pixels whose boundary trends match the color diffusion trends, eliminates false edges caused by lighting or artifacts, and generates a set of structural boundary fitting points that can accurately characterize the expansion activity of the lesion edge.
[0027] S302: Based on the start and end times of the time segment containing the symptom words in the content set of the medical record expression sequence, and using the visual frame number within the corresponding time period as a benchmark, extract the structural position within the corresponding frame in the fitted point set, compare it with the time range of the text segment, and obtain the cross-temporal localization segment value. Based on the spatial localization data of the fitted points set at the structural boundary, and the start and end times of the segments containing symptom words in the corresponding medical record expression sequence content set, the time interval described in the text is mapped to the corresponding dynamic image sequence frame number. During execution, the system retrieves visual frames corresponding to May to August, extracts the spatial coordinates of the fitted point set in frame 4, and calculates the Euclidean distance with the initial coordinates of frame 1, resulting in a pixel displacement of 5.20 pixels. Since this displacement value exceeds the clinical sensitivity benchmark of 5 pixels, the system confirms that the visual geometric expansion and the "range expansion" described in the text are completely consistent in the time dimension. Through this cross-modal spatiotemporal comparison, the system obtains the cross-temporal localization segment values that record the synchronous occurrence of visual feature anomalies and textual descriptions.
[0028] S303: Based on the cross-temporal positioning segment values and the position trajectory of structural transition lines in the image, combined with the regional description information of symptom words in the text, select structural numbers and frame sequence segments with consistent image and text positioning, and integrate them to form a visual language joint expression content group.
[0029] Based on the cross-temporal localization segment values and the positional trajectory of structural transition lines in the image, combined with the directional descriptions of lesion areas in the text using symptom vocabulary, a joint filtering process is performed. During execution, the system uses a spatial logic verification algorithm to determine whether jagged regions in the image with a curvature change rate greater than 0.15 are located within the "left side of the lesion" range described in the text. By filtering out structural numbers and corresponding frame sequence segments with consistent image and text localization, the system precisely semantically anchors abstract textual symptoms to specific image pixel regions. This step integrates visual morphological features with semantic logic in the text, eliminating ambiguities inherent in single-modal analysis and forming a visual language joint expression content group with multi-source feature mutual verification characteristics.
[0030] Please see Figure 5 The steps for obtaining the expression tendency arrangement integrated sequence are as follows: S401: Based on the expression items of each region in the visual language joint expression content group, extract the continuous region of pixel gradient from the color edge in the image, split it into different line segments according to the spatial location, call the disease words in the text segment, determine the start and end positions of each word in the sentence, and obtain the image and text start and end association range by mapping the image edge segment to the text segment words through sequential mapping. Based on the regional expression items in the visual language joint expression content group, the system extracts continuous pixel gradient regions from image edges and uses a local curvature maximum detection algorithm to find the inflection points of curves, thereby splitting complex lesion boundaries into several independent geometric segments based on spatial geometric features. During execution, the system synchronously calls key disease terms such as "irregular," "jagged," and "pigmentation loss" within the text segment. Through word segmentation technology and position anchoring logic, it accurately determines the start and end character indexes of each term in the original long medical record sentence. The system uses a sequential mapping method to establish a logical association between the first edge segment scanned clockwise from a zero-degree angle in the image and the first corresponding attribute word appearing in the reading order in the text segment. Through this one-to-one correspondence processing of image spatial pixel indices and text stream character indices, the system determines the association boundary between specific image edge segments and specific text segment words in the start and end ranges, thus obtaining the image-text start and end association range that can form a closed loop binding between visual physical line segments and textual descriptive semantics.
[0031] S402: Based on the start and end range of the image and text association, extract the path trend of the image edge line segment, calculate the positional order of continuous segments and turning segments in the line segment, call the word order structure of the text symptoms words, determine whether the relationship between adjacent segments in the image path and the flow of language is consistent, and obtain the corresponding expression order. Based on the context of the image and text, the system extracts the directional derivative of the path trend of the image edge segments and uses statistical methods to calculate the spatial arrangement and positional proportion of straight segments and directional turning segments within the line segments. During execution, the system detects 12 alternating positive and negative changes in the sign of the directional derivative within a 100-pixel edge sampling length, determining the turning frequency of that edge segment to be 0.12, and then calls the grammatical structure weights of the text's symptom terms. The system verifies whether the physical features of the image and the descriptive intent of the text point to the same pathological state by judging whether the geometric fluctuation complexity of adjacent segments in the image path has statistical consistency with the semantic strength of morphological adjectives such as "irregular edge height" or "zigzag distribution" in the language flow. By performing this matching and verification of morphological trend and semantic logical flow, the system effectively eliminates false edge fluctuations caused by lighting and shadows or misstatements in medical records, obtaining an objectively quantifiable correspondence between the expression order of the image and text descriptions.
[0032] S403: For cases where the expression order corresponds, the index order is set according to the density of edge line segments in the image area and the distribution order of the disease words, based on the strength of the correspondence between the image and the text. The connection between the regional expression items and the language structure is arranged in sequence to obtain the integrated sequence of expression tendencies.
[0033] For cases involving corresponding expression order, the system combines the dense distribution of edge segments in the image region with the distribution density of disease-related words in the text, and sorts them according to a comprehensive index based on the strength of the connection between the image and text. During execution, the system inputs the calculated semantic feature matching score of 0.85 and the temporal overlap score of 0.90 into a preset weighted fusion model. The calculated comprehensive connection strength of the connection item is 0.87, representing a high confidence level for the image-text association pair. Following the principle of prioritizing all candidate image-text association expressions from high to low connection strength, the visual expressions of the image region and the grammatical connections of the language structure are grouped into a strong association sequence. This step ensures that in subsequent diagnosis, the image-text mutual evidence items with the highest confidence and the most complete evidence chain are prioritized and displayed. Through this sequential logical integration of image-text associations, an integrated sequence of expression tendencies that can guide skin cancer risk identification decisions is obtained.
[0034] Please see Figure 6 The steps to obtain skin cancer risk identification results are as follows: S501: Based on the image starting position in the integrated sequence of expression tendency and the expression content of time symptoms in the text, extract the starting pixel of the corresponding region of the image and the corresponding position of the time words in the paragraph, complete the position matching according to the correspondence between the image sequence and the text order, and obtain the index corresponding to the time image; Based on the image's initial position coordinates (100, 50) in the expression tendency-based integrated sequence and the temporal symptom expressions in the text regarding the onset of the disease course, the system extracts the initial pixel feature points of the corresponding image region and the precise character positions of time-related words in the medical record paragraphs. During execution, the system establishes a correspondence index between the update frequency of the image frame sequence and the disease progression in the text description through linear interpolation algorithms and timeline mapping logic, ensuring that the evolution of each pathological node is traceable and time-aligned in historical images. The system performs deep calibration based on the positional matching accuracy in the spatial sequence distribution of the image and the logical sequence evolution of the text, obtaining a time-image correspondence index that strongly correlates the local features of static images with the evolutionary process of the dynamic disease course narrative. This step realizes the "time stamping" of image regions by the medical record text, laying a data foundation for subsequent analysis of the patterns of disease changes over time.
[0035] S502: Call the image region and disease course expression paragraph associated in the time image corresponding index, identify the extension direction of the color boundary in the image, and compare it with the arrangement order of the disease course content in the language. Based on the relative position of the two, match the boundary region with the word order of the paragraph to obtain the image and text sequence correspondence result. The system retrieves the associated image regions and corresponding disease progression descriptions from the temporal image index, identifies the extension direction vectors of color boundaries in the image, and synchronously compares them with the word order logic of the disease progression in the text description. During execution, the system uses colorimeter analysis technology to identify the dynamic evolution trajectory of color features in the image from an initial light red (RGB values close to 255, 180, 180) to deep black (RGB values close to 20, 20, 20), and determines whether this evolution trend is highly consistent with the textual description of "spot color changing from light to dark and accompanied by diffusion" in terms of temporal relative position. Based on the logical order of their appearance, the system performs precise attribute mapping between the boundary regions and the paragraph word order, dividing the image regions into early stable areas, intermediate diffusion areas, and late infiltrative areas according to the pathological evolution stages. Through this bidirectional verification in the spatiotemporal dimensions, the system obtains image-text sequence correspondence results that can accurately reveal the synchronous evolution of lesions over time and space.
[0036] S503: Based on the image boundary regions and disease expression nodes in the image-text sequence correspondence results, extract the texture change paths in the image and combine them with the concentrated location of the disease content to complete the joint regioning of the image and text, and filter the relevant region content according to the control range to obtain the skin cancer risk identification results.
[0037] Based on the image boundary regions and disease progression expression nodes in the image-text sequence correspondence results, the fractal dimension change paths reflecting the complexity of tissue structure in the image are extracted, and combined with the logical regions with the densest disease progression content description to complete the joint regionization processing of image and text. In the risk identification decision stage, the system extracts the visual abnormality score (0.85), text medical history risk score (0.90), and image-text spatiotemporal mutual verification matching score (0.95) for this region, and calls the preset expert weight allocation parameters for set operation, where the visual weight is set to 0.4, the text weight to 0.3, and the matching weight to 0.3. After weighted operation, the comprehensive risk identification value of this region is obtained as 0.895. Since this value clearly falls into the high-risk judgment range of 0.70 to 1.00, the system automatically judges that this region has an extremely high risk of skin cancer malignancy. Through this series of deep fusion, logical mutual verification, and weighted evaluation of cross-modal features, the system obtains and outputs skin cancer risk identification results containing risk level and lesion location.
[0038] A skin cancer risk assessment system combining visual and language models, the system comprising: The lesion visual acquisition module acquires images of skin lesions, reads the distribution of red, green and blue channels and marks the color concentration areas, detects edge contours and records the edge trend direction, reads the surface texture direction and forms directional labels, sets a sliding window and observes the color changes and edge closure status within the window in turn, screens closed feature areas and performs numbering processing on color jump areas, and generates a sequence of lesion visual changes. The medical record text reconstruction module obtains the medical record text data corresponding to the image, extracts symptom word nodes and marks their position in the sentence segment, extracts time period information and marks the start and end positions, extracts genetic tendency sentence fragments and marks family history keywords, disassembles text paragraphs and deletes duplicate paragraphs, rearranges the semantic order of paragraphs and marks continuous expression relationships, and generates a medical record expression sequence content set. The image-text cross-location module calls the color diffusion area and structural boundary trend in the visual change sequence of the lesion, calls the symptom vocabulary fragments and time description positions in the medical record expression sequence content set, marks the area where the image structure jump line is located and marks the position of the text description sequence segment, performs cross-location of image blocks and text fragments and assigns corresponding numbers, puts them into a unified partition and records the pairing relationship, and generates a visual language joint expression content group. The tendency arrangement integration module calls the image expression items and text expression items of each partition in the visual language joint expression content group, extracts continuous segments of color edges in the image and records their extension paths, extracts the connection mode of symptom words in the text segment and records the order of sentence flow, compares the connection relationship between the image extension path and the order of text flow and marks the consistency status, filters partitions with coherent consistency status and performs front and back sorting processing, merges the sorting results and forms the expression tendency arrangement index, and generates the expression tendency arrangement integration sequence. The risk identification and assessment module, based on the expression tendency arrangement and integration sequence of image start position and text temporal symptom expression content, tracks the color boundary extension area and marks the extension direction and coverage, tracks the disease course description paragraph and marks its start and end time and symptom change description, performs image and text node mapping and merges the image blocks and language structure after nodes to the risk control range, and generates skin cancer risk identification results.
[0039] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method of skin cancer risk assessment combining vision and language models, the method comprising: Comprise the following steps: S1: Collecting skin lesion image color channel distribution, edge trend direction and texture direction, sliding window area observation color change and edge closure, screening color concentration and structure continuity, closing characteristics and color jump area number processing, output lesion visual change sequence; S2: Read the corresponding medical record text of the image, extract the symptom word node, time information and genetic tendency sentence, according to the position information to split the paragraph, delete the repeated expression and rearrange the semantic order, judge whether there is continuous expression relationship according to the word dependent structure, output the expression sequence content set; S3: Based on the color diffusion and structure boundary trend in the lesion visual change sequence, combined with the symptom word and time description in the medical record expression sequence content set, the image structure jump line and the text segment are executed block cross positioning, and the image and text are compared and classified into subarea, and the visual language joint expression content group is output; S4: Call the visual language joint expression content group, extract the image color edge continuous segment and the text disease word connection mode, compare whether the expression start and end range is continuous, process the combination item according to the image change path and language flow direction, and output the arrangement integrated sequence of expression tendency.
2. The method of skin cancer risk assessment of a joint vision and language model according to claim 1, characterized in that: The lesion visual change sequence includes region number label, color jump characteristic value, edge closure degree index and structure continuity parameter, the medical record expression sequence content set includes symptom node list, time period annotation set, genetic tendency segment group and semantic cohesion relationship table, the visual language joint expression content group includes image and text corresponding subarea, structure jump contrast group and symptom time cross node, and the expression tendency arrangement integrated sequence includes image expression path sequence, language structure sorting set and expression coherence mark.
3. The method of skin cancer risk assessment of a joint vision and language model according to claim 1, characterized in that: The acquisition step of the lesion visual change sequence is: S101: Collecting skin lesion image, reading the pixel distribution of red, green and blue color channels in the image, extracting color mutation position according to the standard deviation and mean value of the pixel value in the channel, and obtaining color mutation position coordinate set; S102: Based on the color mutation position coordinate set, set the sliding observation area in the image, judge the edge closure degree according to the edge pixel gray gradient direction and neighborhood distribution relationship in the area, and obtain the closed edge interval value table; S103: Call the closed edge interval value table, combine the color channel value change range and texture direction change trend in the area, screen and number the area meeting the structure closure and color jump condition, and obtain the lesion visual change sequence.
4. The method of skin cancer risk assessment of a joint vision and language model according to claim 1, characterized in that: The acquisition step of the medical record expression sequence content set is: S201: Based on the medical record text data corresponding to the image, extract the symptom word node, time period information and genetic tendency sentence segment, call the position index of various information in the text, split the sentence segment containing related content, and judge whether the position and semantic structure exist repetition, generate the vocabulary and time sentence segment disassembly set; S202: Call the vocabulary and time sentence segment disassembly set, sort the position according to the sentence index and time sequence, identify the word dependent structure between adjacent sentence segments, and arrange the sequence of sentence segments with associated structure, obtain the dependent relationship sentence segment arrangement sequence; S203: According to the arrangement sequence of the dependent relationship sentence segment, the repetition and near-synonymous structure of the adjacent sentence segments are identified, the degree of repetition is judged by combining the position offset information, the deletion or merging operation is performed on the repeated sentence segments, and the medical record expression sequence content set is obtained.
5. The method of skin cancer risk assessment using a combined vision and language model of claim 1, wherein: The acquisition step of the visual language combined expression content group is: S301: Call the color diffusion part and structure boundary trend in the lesion visual change sequence, map the color value gradient change and the boundary line trend, compare the difference degree of the boundary angle change and the color diffusion edge, extract the points consistent with the structure boundary trend in the fitting point set, and generate the structure boundary fitting point set; S302: According to the time start and end position of the symptom vocabulary in the structure boundary fitting point set and the medical record expression sequence content set, the visual frame number in the corresponding time period is taken as a reference, the structure position in the corresponding frame in the fitting point set is extracted, and the time range of the text segment is compared to obtain the cross timing positioning section value; S303: According to the cross timing positioning section value and the position trajectory of the structure jump line in the image, combined with the area description information of the symptom vocabulary in the text, the structure number and frame sequence number segment consistent with the text positioning are screened and integrated to form the visual language combined expression content group.
6. The method of skin cancer risk assessment of a joint vision and language model according to claim 1, wherein: The acquisition step of the expression tendency arrangement integrated sequence is: S401: Based on each area expression item in the visual language combined expression content group, the pixel gradient continuous area of the color edge in the image is extracted, which is divided into different line segments according to the spatial position, the start and end positions of each word in the sentence are determined by calling the disease words in the text segment, the image edge segment and the text segment word are corresponded by the sequential mapping method, and the start and end associated range of the image and the text is obtained; S402: According to the start and end associated range of the image and the text, the path trend of the image edge line segment is extracted, the position sequence of the continuous section and the turning section in the line segment is calculated, the language sequence structure of the text disease word is called, whether the front and rear relationship of the adjacent segments in the image path and the language flow direction is consistent is judged, and the expression order corresponding condition is obtained; S403: According to the expression order corresponding condition, combined with the dense relationship of the edge line segment in the image area and the distribution order of the disease words, the connection mode between the area expression item and the language structure is arranged in sequence according to the corresponding connection strength between the image and the text, and the expression tendency arrangement integrated sequence is obtained.
7. The method of skin cancer risk assessment of a joint vision and language model according to claim 1, wherein, The method further comprises the S5 step: S5: Call the image start position and the text time symptom expression in the expression tendency arrangement integrated sequence, track the color boundary extension area, compare the disease course paragraph in the language expression, and after comparing the continuous content joints of the two types of sources, map the image block and the language content to the risk area, and output the skin cancer risk identification result; The skin cancer risk identification result includes a risk image block set, a disease course description paragraph group, and a text joint comparison table.
8. The method of skin cancer risk assessment of a joint vision and language model according to claim 7, characterized in that: The acquisition step of the skin cancer risk identification result is: S501: based on the expression tendency arrangement integrated sequence, the starting position of the image and the expression content of the time symptom in the text are arranged, the starting pixel point of the image corresponding area and the corresponding position of the time word in the paragraph are extracted, the position matching is completed according to the corresponding relationship in the image sequence and the text sequence, the time image corresponding index is obtained; S502: calling the image area and the course expression paragraph associated in the time image corresponding index, identifying the extension direction of the color boundary in the image, and comparing the arrangement order of the course content in the language, the boundary area and the word sequence are matched according to the relative position of the two, and the image-text sequence corresponding result is obtained; S503: according to the image boundary area and the course expression node in the image-text sequence corresponding result, the texture change path in the image is extracted and combined with the concentrated position of the course content, the joint region of the image and the text is completed, and the related area content is screened according to the comparison range, and the skin cancer risk identification result is obtained.
9. A skin cancer risk assessment system that combines a vision and language model, characterized by, The system is used for the skin cancer risk assessment method of the joint visual and language model in any one of claims 1-8, and the system comprises: a lesion visual acquisition module, which acquires a skin lesion image, reads a red-green-blue channel distribution and labels a color concentrated area, detects an edge contour and records an edge trend direction, reads a surface texture direction and forms a direction label, sets a sliding window and sequentially observes color change and edge closure state in the window, screens a closed feature area and performs numbering processing on a color jump area, and generates a lesion visual change sequence; a medical record text reconstruction module, which acquires medical record text data corresponding to the image, extracts symptom vocabulary nodes and labels their sentence segment positions, extracts time period information and labels start and end positions, extracts genetic tendency sentence segments and labels family history pointing words, disassembles text paragraphs and deletes duplicate paragraphs, rearranges paragraph semantic order and labels continuous expression relationship, and generates a medical record expression sequence content set; a cross positioning module, which calls the color diffusion area and the structure boundary trend in the lesion visual change sequence, calls the symptom vocabulary segment and the time description position in the medical record expression sequence content set, labels the area where the image structure jump line is located and labels the position where the text description sequence is located, performs cross positioning of the image block and the text segment and performs numbering corresponding, is classified into a unified partition and records the pairing relationship, and generates a visual language joint expression content group; a tendency arrangement integration module, which calls the image expression item and the text expression item in each partition of the visual language joint expression content group, extracts color edge continuous segments in the image and records their extension paths, extracts disease word connection mode in the text segment and records sentence flow direction order, compares the front and back connection relationship of the image extension path and the text flow direction order and labels the consistency state, selects the partitions with consistent consistency state and performs front and back sorting processing, combines the sorting results and forms an expression tendency arrangement index, and generates an expression tendency arrangement integrated sequence. The risk identification and evaluation module integrates the image starting position and the text time symptom expression content in the expression tendency arrangement set, tracks the color boundary extension area and labels the extension direction and coverage range, tracks the course description paragraph and labels its time start and end and symptom change description, performs image joint and text joint mapping and merges the image blocks and language structures after joint to the risk control range, and generates a skin cancer risk identification result.
Citation Information
Patent Citations
Electronic clinical medical assessment method and system for ophthalmology diagnosis and treatment scheme
CN120564937A
Document retrieval method, system and equipment and medium
CN120632057A
Radiotherapy target volume identification method based on large language model
CN121033461A
Burn scar hyperplasia risk assessment method combining semantic segmentation and texture feature analysis
CN121392444A
Content based image retrieval for lesion analysis
US20200380675A1
Cited By
Skin cancer multi-classification modeling method based on data enhancement and model optimization
CN121837798A