Multi-modal resume layout adaptive analysis method and system based on large model
Through a multimodal resume layout adaptive parsing method that combines a large model with edge density and language recognition, the problem of poor adaptability of multimodal resumes is solved, and efficient resume content optimization and sorting are achieved. It is suitable for mixed text and image layout and resume scenarios with diverse styles in multilingual environments.
Patent Information
- Application Number
- CN202510871351.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing resume parsing systems have poor adaptability and low accuracy when dealing with multimodal and multi-style resumes, making it difficult to meet the needs of corporate HR for high-quality candidate screening.
A multimodal resume layout adaptive parsing method based on a large model is adopted. The image format is unified through edge density calculation and language recognition. Functional areas are identified by combining OCR recognition, visual and semantic feature fusion, and content sorting is optimized through vector matching and clustering statistics.
It achieves accurate analysis and content optimization of multimodal resumes, improves the accuracy and adaptability of information extraction, and generates a personalized resume structure with clear structure and highlighted key points.
Smart Images

Figure CN120763397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of adaptive parsing technology, and more specifically, to a method and system for adaptive parsing of multimodal resume layout based on a large model. Background Art
[0002] With the development of information technology, more and more job seekers are choosing to submit their resumes online electronically. Due to the wide variety of resume templates and content layout styles used by different users, traditional resume parsing methods based on templates or fixed rules often suffer from poor adaptability and low accuracy.
[0003] Furthermore, resumes contain diverse content in various forms, such as text, images, and tables, constituting typical multimodal data. Existing parsing systems are prone to problems such as failure to identify functional areas, incomplete information extraction, and illogical sorting when faced with resumes containing a mixture of text and images, in different languages, or with low-definition images. This makes it difficult for HR to meet the needs of high-quality candidate screening.
[0004] Therefore, there is an urgent need for a resume parsing system with higher intelligence and adaptability to achieve accurate parsing and content optimization of multimodal and multi-style resumes. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a multimodal resume layout adaptive parsing method and system based on a large model to solve the problems raised in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions: The multimodal resume layout adaptive parsing method based on a large model includes the following steps: first, the resumes in various formats uploaded by users are uniformly converted into a standard image format, edge density calculation and language recognition are performed, and the image enhancement coefficient is calculated based on the edge density and language set. The calculated image enhancement coefficient value determines whether to perform image enhancement; Automatically classify resume paragraphs by functional area, then find the TOP-K similar resumes through vector matching and perform clustering statistics. Combine HR attention and the difference in word count of each functional area to calculate the weight coefficient of each area in the parsed document. Then, sort the resume content and optimize its length based on the weight coefficient of each area in the parsed document.
[0007] In a preferred embodiment, the edge density is calculated using the following formula: ; Where: |E| represents the number of edge pixels; W, H are the width and height of the image respectively.
[0008] In a preferred embodiment, the language recognition is based on a pre-trained language detection model, analyzing the language distribution in the image I and outputting a language set L: ; The language set L reflects the number of language types in the image.
[0009] In a preferred embodiment, the edge density r and the language set L are determined, normalized first, and then the image enhancement coefficient is calculated by weighted summation.
[0010] In a preferred embodiment, an image enhancement coefficient is determined. When the image enhancement coefficient is greater than a system threshold, image enhancement is not required; conversely, when the image enhancement coefficient is less than the system threshold, image enhancement is required.
[0011] In a preferred embodiment, if an image is determined to require enhancement, super-resolution reconstruction is first performed by introducing the ESRGAN-Doc model, which is based on the ESRGAN architecture and incorporates a row-column attention mechanism. A cross-resolution consistency learning mechanism is used to compare the OCR output results of the original image and the enhanced image using KL divergence to achieve recognition consistency optimization. OCR recognition uses a dual-channel structure: one channel is a CTC-Transformer model that supports multi-language recognition, and the other channel is optimized in conjunction with a language-specific dictionary. When there is a difference between the outputs of the two channels, a language model is introduced to calculate the perplexity and select the text with the higher language naturalness as the final output. Subsequently, the recognized text is formatted and invalid characters are removed. The paragraph boundaries are reconstructed based on the OCR position information to complete the paragraph structuring process. LayoutLMv3 is introduced to extract visual features, and combined with a large language model to extract semantic features, a unified multimodal representation is constructed through the cross-attention mechanism. Considering the problem of OCR boundary errors, the Diffusion Layout Encoder is further introduced to simulate layout disturbances through the diffusion mechanism and achieve reverse position reconstruction. Finally, the corrected paragraph position box and saliency heat map are output.
[0012] In a preferred embodiment, based on the output saliency heat map and multimodal semantic vector, rule guidance and BERT multi-label classification model are combined to predict functional area labels for paragraphs and identify different functional areas.
[0013] In a preferred embodiment, after identification, the system extracts key feature vectors from the user's resume, including: education, job objectives, and keyword embeddings; using this as the query vector, it performs cosine similarity calculation in the constructed resume vector library and selects the top-k similar samples: ;in Feature vector extracted from the user's resume, used for similar resume retrieval; Represents the feature vector in the reference resume; is the cosine similarity between the user and the reference resume; is the set of top K reference samples with the highest similarity to the user's resume; The reference resume is a historically submitted resume in the resume vector library.
[0014] In a preferred embodiment, the content of each functional area in these similar resumes is then clustered to obtain the main expression method of each area; based on this, the HR attention ranking can be formed: ;in The ranking of functional areas sorted by HR attention is the HR attention ranking; is the frequency of the region appearing in the reference sample; For the functional area r in the user's resume, count the actual number of words , calculate the difference between the number of words and the historical average: ; Represents the difference in word count from the historical average, the historical average Obtained from the average of historical resume submissions in the resume vector library; Calculate the difference between the word count of each functional area and the historical mean in the user's resume, calculate the average of the word count differences of each functional area and the historical mean, and calculate the difference between the word count differences of each functional area and the historical mean and the average of the word count differences of each functional area and the historical mean; the difference between the word count differences of each functional area and the historical mean and the average of the word count differences of each functional area and the historical mean is simplified to express as the word count difference of each functional area; Determine the HR attention ranking and the difference in word count of each functional area; first perform normalization, and then calculate the weight coefficient of each area in the parsed document through weighted summation; sort and combine the content of each functional area according to the value of the weight coefficient of each area in the parsed document; the functional area with a larger regional weight coefficient has more words and is ranked higher, and the functional area with a smaller regional weight coefficient has fewer words and is ranked lower; based on the user resume, reduce the content of the functional area with a small regional weight coefficient.
[0015] In a preferred embodiment, the system includes the following modules: input preprocessing module, structured analysis module, functional area identification module, content optimization and recommendation module; The input preprocessing module is used for receiving a user uploaded resume file in multiple formats, and uniformly converting the resume file into a standard image format according to a corresponding conversion mode of different file types; the converted image is input into a quality detection and language recognition process, an edge detection algorithm is used to calculate edge density, a language recognition model is used to judge the language type, and an image enhancement coefficient is calculated according to the edge density and the language set to determine whether image enhancement is performed; The structured analysis module receives the initial text result obtained by OCR recognition, first performs format cleaning, then reconstructs the paragraph boundary according to the OCR position information, and forms a paragraph area through character position aggregation; subsequently, the system introduces LayoutLMv3 to extract the visual feature vector of the paragraph image, and calls a large language model to generate a corresponding semantic feature representation, fuses the two features through a cross attention mechanism, and constructs a unified multi-modal representation; on this basis, the system further adopts Diffusion Layout Encoder to reconstruct the layout position of the paragraph position box, generates more accurate paragraph coordinates through a diffusion denoising mechanism, and outputs a saliency heat map as a reference for subsequent functional area recognition; The functional area recognition module is based on the paragraph position and multi-modal semantic embedding result output by the structured analysis module, combines the saliency heat map, and uses rule guidance and a BERT multi-label classification model to label the paragraph to realize automatic recognition and classification of the key area of the resume; The content optimization and recommendation module extracts a key feature vector from the user's resume on the basis of the functional area recognition, searches for TOP-K similar samples in the constructed resume vector library through cosine similarity matching, and performs clustering analysis on the content of each functional area in the reference resume, and calculates the average number of words, the standard deviation and the frequency of occurrence of each area, and then constructs an HR attention ranking; combined with the deviation of the word number of each functional area in the user's resume from the average value of the reference sample, the weight coefficient of each area is calculated, and the system sorts and adjusts the length of the resume content according to the weight coefficient of each area.
[0016] The technical effects and advantages of the present application are as follows: The multi-modal resume layout adaptive analysis method and system based on a large model provided by the present application have good adaptability and intelligence, can compatible multiple resume formats and uniformly convert them into images, calculate an image enhancement coefficient through edge density and language type, intelligently judge whether to perform image enhancement, and effectively improve the OCR recognition accuracy.
[0017] The system integrates a visual model with a large language model, constructing a unified multimodal representation through a cross-attention mechanism. It also introduces a Diffusion Layout Encoder to accurately correct paragraph layout and enhance structured parsing capabilities. Furthermore, it combines saliency heatmaps with the BERT multi-label classification model to accurately identify the functional areas of resumes.
[0018] The system also matches historical samples based on user resume features, calculates regional weights based on HR attention and word count deviation, and intelligently sorts and optimizes content length, ultimately generating a personalized resume structure with clear structure and focused emphasis. This method supports resume parsing in multilingual environments and is suitable for complex resume scenarios with mixed text and images and diverse styles. It significantly improves information extraction accuracy and content recommendation effectiveness, and has broad practical value and industrial application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 Schematic diagram of the process of the adaptive parsing method for multimodal resume layout based on a large model of the present invention; Figure 2 This is a flowchart of the image enhancement process of the present invention; Figure 3 This is a structural diagram of the multimodal resume layout adaptive parsing system based on a large model of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] Example 1, the present invention is based on a multi-modal resume layout adaptive parsing method of a large model, such as Figure 1 As shown, the following steps are included: First, the various resume formats uploaded by users are converted into a standard image format, edge density calculation and language recognition are performed, and the image enhancement coefficient is calculated based on the edge density and language set. The calculated image enhancement coefficient value determines whether to perform image enhancement; The recognition phase uses a dual-channel architecture of a CTC-Transformer model and a language-specific dictionary, and optimizes the final output text through language model perplexity evaluation. The OCR results are then formatted and paragraphs reconstructed. LayoutLMv3 and a large language model are combined to extract visual and semantic features, which are then fused into a multimodal representation using a cross-attention mechanism. A Diffusion Layout Encoder is also introduced to correct paragraph position and generate saliency heatmaps. Based on rule guidance and BERT multi-label classification, resume paragraphs are automatically classified into functional areas. Then, through vector matching, the TOP-K similar resumes are found and clustered and counted. The weight coefficient of each area in the parsed document is calculated by combining HR attention and the difference in the number of words in each functional area. The resume content is sorted and the length is optimized according to the weight coefficient of each area in the parsed document.
[0022] Specifically, the resume file uploaded by the user may be in the following formats: image type: JPG, PNG; document type: PDF, DOCX; text type: TXT.
[0023] First select the appropriate converter based on the file type , uniformly converted into a standard image format (such as PNG): Image conversion formula: ; Where: D represents the original document; I is the converted standard image input.
[0024] Furthermore, PDF and Word documents will be rendered as images through page screenshots or PDF renderers; TXT text will be synthesized into images using standard typesetting templates.
[0025] The converted image will enter the quality inspection and language recognition process. The specific steps are as follows: Image clarity is obtained by edge detection algorithm (such as Canny) to obtain edge map E, and edge density r is calculated: Edge density formula: Where |E| represents the number of edge pixels; W and H represent the image width and height, respectively. This value reflects the density and clarity of text in the image.
[0026] Based on the pre-trained language detection model (such as fastText, multilingual OCR classification head), analyze the language distribution in image I and output the language set L: ; The language set L reflects the number of languages in the image.
[0027] Determine the edge density r and the language set L, first perform normalization processing, and then calculate the image enhancement coefficient using the weighted sum formula. The specific formula is as follows: X = α * r - β * L; where X represents the image enhancement coefficient; the greater the edge density, the greater X, and vice versa; the larger the language set, the smaller X, and vice versa; α and β are the weight coefficients of the edge density r and the language set L, respectively, and can be set as needed according to actual needs.
[0028] Whether to perform image enhancement is determined based on the value of the image enhancement coefficient. Specifically, when the image enhancement coefficient is greater than the system threshold, it means that the text density and clarity in the image are high and the number of language types in the image is small, and image enhancement is not required. Conversely, when the image enhancement coefficient is less than the system threshold, it means that the text density and clarity in the image are low and the number of language types in the image is large, and image enhancement is required.
[0029] If the image is determined to be enhanced, such as Figure 2 As shown, the ESRGAN-Doc module is used for image super-resolution reconstruction. This model is based on the original ESRGAN and incorporates the row / column attention mechanism to enhance the modeling ability of text layout. The output of the enhanced image is: ; This is a high-resolution image enhanced by ESRGAN; It is a super-resolution reconstruction model used to improve image clarity.
[0030] The enhanced image will be used to improve OCR recognition accuracy. To further enhance recognition stability, a cross-resolution consistency learning mechanism is introduced, using KL divergence to measure the difference between the original image and the enhanced image OCR output: ;in The probability distribution of the OCR recognition output for low-resolution images; The probability distribution of the OCR recognition output for high-resolution images; It is the KL divergence loss between low and high resolution outputs, used for consistency optimization.
[0031] OCR recognition uses a dual-channel architecture. The first channel is a CTC-Transformer model with multilingual recognition capabilities in Chinese, English, and French. The second channel uses language-specific dictionaries (such as a Chinese name dictionary and an English job title dictionary) to optimize output. If the edit distance between the two channel recognition results is greater than the edit distance, a language model (such as GPT) is used to calculate the perplexity: ;in, Output text for the final selected OCR (with the lowest perplexity); Represents the candidate text outputs from the two OCR channels. Perplexity(s) is the perplexity of the language model for text s. Lower values indicate more fluent and natural language. The text with the lower perplexity is ultimately selected as the OCR output.
[0032] Furthermore, due to the presence of noise and paragraph misalignment in the OCR text results, a format cleaning and reconstruction module is required. First, the text is cleaned to remove invalid characters, special symbols, and garbled characters. Then, a paragraph area box is constructed based on the position information returned by OCR. A paragraph is represented by all the characters within the position box: ;in The bounding box (position coordinates) of the i-th paragraph identified by OCR; is the jth character in paragraph i; is the set of characters that make up paragraph i (aggregated by position).
[0033] The system unifies the line spacing and normalizes the text for each paragraph, and outputs a structured intermediate representation: To achieve the fusion of semantics and visual layout, the system introduces LayoutLMv3 to extract visual feature vectors. , and combined with large language models (such as ChatGPT) to extract semantic representations , fused through the cross-attention mechanism: ; H represents the multimodal representation that integrates vision and text, and the embedding result after cross attention.
[0034] Furthermore, due to the deviation of the OCR output boundary, the Diffusion Layout Encoder module is required to reconstruct the position. The diffusion mechanism gradually adds Gaussian noise: ;in Represents the original real layout coordinates (the paragraph box in the ideal state); represents the layout representation at diffusion step t (noisy version); represents the diffusion noise attenuation coefficient of the t-th step, controlling the degree of information retention; represents standard Gaussian noise, simulating layout disturbances.
[0035] Through the diffusion model To achieve the original position Reverse reconstruction: ;in is the denoising function of the diffusion model with parameter θ, used to predict ;in This is a language condition vector (e.g., a one-hot encoding of the language category) that guides the layout generation process. Outputs include: corrected paragraph location boxes and saliency heatmaps for functional area identification.
[0036] This stage aims to identify key functional areas in the resume and output personalized optimization suggestions through content clustering and statistics.
[0037] First, based on the saliency heat map and multimodal semantic vector output in the previous stage, the system combines rule guidance with the BERT multi-label classification model to predict functional area labels for the paragraphs, and identifies areas such as "basic information", "educational experience", "project experience", "skills and expertise", and "awards and honors".
[0038] After identification, the system extracts key feature vectors from the user's resume, including: education (one-hot), job objectives (TF-IDF vector), keyword embedding, etc. Using this as the query vector, it performs cosine similarity calculation on the established resume vector library and selects the top-k similar samples: ;in Feature vector extracted from the user's resume, used for similar resume retrieval (such as education level and keywords); Represents the feature vector in the reference resume; is the cosine similarity between the user and the reference resume; is the set of top K reference samples with the highest similarity to the user's resume.
[0039] The reference resume is a historically submitted resume in the resume vector library.
[0040] Then, we perform K-Means or DBSCAN clustering on the content of each functional area in these similar resumes to obtain the main expression methods of each area, and calculate the following parameters: average word count, standard deviation, and frequency of occurrence. Based on this, we can form the HR attention ranking: ;in The ranking of functional areas sorted by HR attention is the HR attention ranking; The frequency of this area appearing in the reference sample reflects HR attention. The higher the HR attention ranking, the greater the HR attention to this functional area.
[0041] For the functional area r in the user's resume, count the actual number of words , calculate the difference between the number of words and the historical average: ; Represents the difference in word count from the historical average, the historical average Obtained from the average of historically submitted resumes in the resume vector library.
[0042] The difference between the number of words in each functional area in the user's resume and the historical average is calculated, the average of the difference between the number of words in each functional area and the historical average is calculated, and the difference between the difference between the number of words in each functional area and the historical average and the average of the difference between the number of words in each functional area and the historical average is calculated; avoid all functional areas of the user's resume having more words; affect the judgment. The difference between the difference between the number of words in each functional area and the historical average and the average of the difference between the number of words in each functional area and the historical average is simplified as the difference between the number of words in each functional area.
[0043] The greater the difference between the number of words in a certain functional area of the user's resume, the more the user wants to draw the attention of HR to the content of this area, and the content proportion of this area should be appropriately adjusted.
[0044] Determine the HR attention ranking and the number of words in each functional area; first, normalize the processing, and then calculate the weight coefficient of each area in the parsed document by weighted summation; the specific formula is as follows: ; wherein represents the weight coefficient of functional area r in the parsed document; c(r) represents the number of words in functional area r, wherein the greater the number of words in functional area r, the greater the weight coefficient of functional area r in the parsed document, and vice versa; Rank(r) represents the HR attention ranking of functional area r; wherein the greater the HR attention ranking of functional area r, the greater the weight coefficient of functional area r in the parsed document, and vice versa; η and λ are the weight coefficients of the HR attention ranking and the number of words in each functional area, respectively, which are set as needed.
[0045] Determine the weight coefficient of each area in the parsed document, and sort and combine the content of each functional area according to the numerical value of the weight coefficient of each area in the parsed document; wherein the greater the weight coefficient of the functional area, the more the number of words and the more the arrangement in front, and for the functional area with smaller weight coefficient, the less the number of words and the more the arrangement in the back; based on the user's resume, the content of the functional area with small weight coefficient is reduced.
[0046] Embodiment 2, the design of the multi-modal resume layout adaptive analysis system based on the large model of the present application is based on the method in embodiment 1, as shown in Figure 3 includes the following modules: input preprocessing module, structured analysis module, functional area identification module, content optimization and recommendation module; The input preprocessing module is used to receive resume files in various formats uploaded by users, including image, document and text types, and unify them into standard image formats using corresponding conversion methods according to different file types. PDF and Word files are generated by rendering, and TXT texts are synthesized into images through standard typesetting templates. The converted image will be input into the quality detection and language recognition process, and the edge density will be calculated using the edge detection algorithm. The language type will be determined by the language recognition model, and the image enhancement coefficient will be further calculated to determine whether image enhancement is required. When enhancement is required, the ESRGAN-Doc super-resolution reconstruction model is called to improve the image quality. The enhanced image and the original image are respectively input into the OCR module, and the CTC-Transformer and language dictionary dual-channel structure is used to recognize the text, and the optimal text output result is selected through the perplexity evaluation mechanism.
[0047] The structured parsing module is used to receive the initial text results obtained by OCR recognition, first perform format cleaning, remove garbled characters and abnormal characters, then reconstruct the paragraph boundaries based on the OCR position information, and form paragraph areas through character position aggregation. Subsequently, the system introduces LayoutLMv3 to extract the visual feature vector of the paragraph image, and calls the large language model to generate the corresponding semantic feature representation. The two features are fused through the cross-attention mechanism to construct a unified multimodal representation. On this basis, in order to correct the OCR boundary error, the system further uses Diffusion Layout Encoder to reconstruct the layout position of the paragraph position box, uses the diffusion denoising mechanism to generate more accurate paragraph coordinates, and outputs the significance heat map as a reference for subsequent functional area identification.
[0048] The functional area identification module, based on the paragraph positions and multimodal semantic embedding results output by the structured parsing module, combines this with a saliency heatmap and employs rule-based guidance and the BERT multi-label classification model to label paragraphs with functional areas. This allows for the automatic identification and classification of key resume areas such as "basic information," "educational experience," "project experience," and "skills and expertise." Upon completion of this phase, the resume text is segmented by semantic categories and mapped to a visual layout, achieving a clear and visually organized functional structure.
[0049] Based on functional area identification, the content optimization and recommendation module extracts key feature vectors from the user's resume, retrieves TOP-K similar samples in the constructed resume vector library through cosine similarity matching, and performs cluster analysis on the content of each functional area in the reference resume. The average number of words, standard deviation and frequency of occurrence of each area are counted to construct an HR attention ranking. The regional weight coefficient is calculated based on the deviation between the number of words in each functional area in the user's resume and the mean of the reference sample. The HR attention and word count difference are combined through a weighted normalization formula to output the final functional area weight. The system sorts and adjusts the length of the resume content according to the weight, recommends compressing the content of areas with small weights, and highlights those with large weights. Finally, it generates typesetting optimization suggestions to achieve intelligent reconstruction of resume content.
[0050] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0051] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0052] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0053] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0054] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A multimodal resume layout adaptive parsing method based on a large model, characterized by: The following steps are involved: First, the various resume formats uploaded by users are converted into a standard image format, edge density calculation and language recognition are performed, and the image enhancement coefficient is calculated based on the edge density and language set. The calculated image enhancement coefficient value determines whether to perform image enhancement; Automatically classify resume paragraphs by functional area, then find the TOP-K similar resumes through vector matching and perform clustering statistics. Combine HR attention and the difference in word count of each functional area to calculate the weight coefficient of each area in the parsed document. Then, sort the resume content and optimize its length based on the weight coefficient of each area in the parsed document.
2. The large-model-based multimodal resume layout adaptive parsing method according to claim 1 is characterized by: The edge density calculation formula is as follows: ; Where: |E| represents the number of edge pixels; W, H are the width and height of the image respectively.
3. The large-model-based multimodal resume layout adaptive parsing method according to claim 1 is characterized by: The language recognition, based on the pre-trained language detection model, analyzes the language distribution in the image I and outputs the language set L: ; The language set L reflects the number of language types in the image.
4. The large-model-based multimodal resume layout adaptive parsing method according to claim 3 is characterized by: Determine the edge density r and the language set L, first perform normalization, and then calculate the image enhancement coefficient through weighted summation.
5. The large model-based multimodal resume layout adaptive parsing method according to claim 4 is characterized by: Determine the image enhancement coefficient. When the image enhancement coefficient is greater than the system threshold, image enhancement is not required; conversely, when the image enhancement coefficient is less than the system threshold, image enhancement is required.
6. The large-model-based multimodal resume layout adaptive parsing method according to claim 5, characterized in that: If an image is determined to require enhancement, super-resolution reconstruction is first performed by introducing the ESRGAN-Doc model, which is based on the ESRGAN architecture and incorporates a row-column attention mechanism. A cross-resolution consistency learning mechanism is used to compare the OCR output results of the original and enhanced images using KL divergence to optimize recognition consistency. OCR recognition uses a dual-channel structure: one channel is a CTC-Transformer model that supports multilingual recognition, and the other channel is optimized using a language-specific dictionary. When there is a discrepancy between the two channel outputs, a language model is introduced to calculate the perplexity and select the text with the higher linguistic naturalness as the final output. Subsequently, the recognized text is formatted and invalid characters are removed, and paragraph boundaries are reconstructed in combination with OCR position information to complete paragraph structuring. LayoutLMv3 is introduced to extract visual features, and combined with a large language model to extract semantic features, a unified multimodal representation is constructed through a cross-attention mechanism. Considering the OCR boundary error problem, DiffusionLayout Encoder is further introduced to simulate layout disturbances and achieve reverse position reconstruction through a diffusion mechanism, and finally output the corrected paragraph position box and saliency heat map.
7. The large model-based multimodal resume layout adaptive parsing method according to claim 6, characterized in that: Based on the output saliency heatmap and multimodal semantic vector, combined with rule guidance and the BERT multi-label classification model, functional area labels are predicted for paragraphs to identify different functional areas.
8. The large model-based multimodal resume layout adaptive parsing method according to claim 7, characterized in that: After identification, the system extracts key feature vectors from the user's resume, including education, job objectives, and keyword embeddings. Using these as query vectors, it performs cosine similarity calculations on the established resume vector library and selects the top-k similar samples: ;in Feature vector extracted from the user's resume, used for similar resume retrieval; Represents the feature vector in the reference resume; is the cosine similarity between the user and the reference resume; is the set of top K reference samples with the highest similarity to the user's resume; The reference resume is a historically submitted resume in the resume vector library.
9. The large model-based multimodal resume layout adaptive parsing method according to claim 8, characterized in that: Then, we clustered the content of each functional area in these similar resumes to obtain the main expression methods of each area; based on this, we formed the HR attention ranking: ;in The ranking of functional areas sorted by HR attention is the HR attention ranking; is the frequency of the region appearing in the reference sample; For the functional area r in the user's resume, count the actual number of words , calculate the difference between the number of words and the historical average: ; Represents the difference in word count from the historical average, the historical average Obtained from the average of historical resume submissions in the resume vector library; Calculate the difference between the word count of each functional area and the historical mean in the user's resume, calculate the average of the word count differences of each functional area and the historical mean, and calculate the difference between the word count differences of each functional area and the historical mean and the average of the word count differences of each functional area and the historical mean; the difference between the word count differences of each functional area and the historical mean and the average of the word count differences of each functional area and the historical mean is simplified to express as the word count difference of each functional area; Determine the HR attention ranking and the difference in word count of each functional area; first perform normalization, and then calculate the weight coefficient of each area in the parsed document through weighted summation; sort and combine the content of each functional area according to the value of the weight coefficient of each area in the parsed document; the functional area with a larger regional weight coefficient has more words and is ranked higher, and the functional area with a smaller regional weight coefficient has fewer words and is ranked lower; based on the user resume, reduce the content of the functional area with a small regional weight coefficient.
10. A multimodal resume layout adaptive parsing system based on a large model, characterized by: The parsing system is based on the method according to any one of claims 1 to 10, and comprises the following modules: an input preprocessing module, a structured parsing module, a functional area identification module, and a content optimization and recommendation module; The input preprocessing module is used to receive resume files in various formats uploaded by users and convert them into a standard image format using corresponding conversion methods according to different file types. The converted images will be input into the quality inspection and language recognition process, where edge density is calculated using an edge detection algorithm, and the language type is determined through a language recognition model. The image enhancement coefficient is further calculated based on the edge density and language set to determine whether image enhancement should be performed. After receiving the initial text results from OCR recognition, the structured parsing module first performs formatting cleanup, then reconstructs paragraph boundaries based on OCR position information, and forms paragraph regions by aggregating character positions. The system then uses LayoutLMv3 to extract visual feature vectors from paragraph images and calls a large language model to generate corresponding semantic feature representations. The two features are then fused through a cross-attention mechanism to construct a unified multimodal representation. Furthermore, the system uses a Diffusion Layout Encoder to reconstruct the layout position of the paragraph position box, using a diffusion denoising mechanism to generate more accurate paragraph coordinates. Furthermore, it outputs a saliency heat map as a reference for subsequent functional region identification. The functional area recognition module uses the paragraph position and multimodal semantic embedding results output by the structured parsing module, combined with the saliency heat map, to label the paragraphs with functional areas using rule guidance and the BERT multi-label classification model, thereby automatically identifying and classifying key areas of the resume. Based on functional area identification, the content optimization and recommendation module extracts key feature vectors from user resumes, retrieves TOP-K similar samples in the established resume vector library through cosine similarity matching, and performs cluster analysis on the content of each functional area in the reference resumes. The module calculates the average word count, standard deviation, and frequency of occurrence of each area, and then constructs an HR attention ranking. Based on the deviation between the number of words in each functional area of the user's resume and the mean of the reference sample, the weight coefficient of each area is calculated. The system sorts the resume content and adjusts the length according to the weight coefficient of each area.
Citation Information
Cited By
Intelligent screening method, device and system for intelligent personnel platform
CN121032454A