Post adaptation degree evaluation method and system based on multi-modal resume analysis

CN122529677APending Publication Date: 2026-08-07FUJIAN JUNNUO SCI & TECH ACHIEVEMENTS TRANSFORMATION SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN JUNNUO SCI & TECH ACHIEVEMENTS TRANSFORMATION SERVICE CO LTD
Filing Date
2026-06-30
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于多模态简历解析的岗位适配度评估方法及系统,其主要目的在于解决基于多模态简历解析的岗位适配度评估时准确率较低的问题

Benefits of technology

[0016]1.本技术通过将待评估简历文件分离为简历文本流和简历图像区域集,并将图像区域集中的像素数据转化为图像语义凭据,再基于版面位置关系将图像语义凭据关联至对应的文本条目,构建了多模态对齐简历档案,实现了文本信息与图像信息的深度融合和版面级精准挂载,使得原本孤立于简历中的证书、照片等视觉内容能够作为对应技能或经历的附属证据直接参与岗位适配度评估。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122529677A_ABST
    Figure CN122529677A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of talent recruitment, and discloses a post adaptation degree evaluation method and system based on multi-modal resume analysis, the method comprising: obtaining a post description text and a resume file to be evaluated, separating the resume into a text stream and an image region set; information tuples in the text stream are taken as a resume semantic archive, and pixel data of the image region are converted into image semantic credentials; based on the layout position relationship between the text block and the image region, the image semantic credentials are associated to the corresponding text entry to obtain a multi-modal aligned resume archive; the skills and qualifications fields of the post description text are collected into a post requirement set; the demand is taken as an index to position the matching field and the verification record to obtain an evidence group; according to the mutual verification strength of the hit entry and the supporting text, a post adaptation evaluation result is generated; the present application can improve the accuracy of post adaptation degree evaluation based on multi-modal resume analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of talent recruitment technology, and in particular to a method and system for evaluating job suitability based on multimodal resume parsing. Background Technology

[0002] Existing resume parsing technologies have significant shortcomings when processing multimodal content. Most methods only extract and match plain text fields in resumes, completely ignoring the visual information contained in image areas such as certificate photos, skill badges, and personal portraits. This results in a large amount of evidence that can be used to verify job suitability being discarded, severely limiting the comprehensiveness and credibility of the evaluation results.

[0003] In existing job suitability assessment methods, even if a few technologies attempt to introduce image information, there is a lack of effective means to align the semantics of images with the corresponding text entries at the layout level. Usually, images are processed independently or simply stacked at the end of the text, which fails to establish the attribution relationship between the image and the text describing its content. This makes it impossible to accurately attach image evidence to the corresponding skills or experience fields, resulting in broken evidence chains, mismatched positioning, and low accuracy of the final assessment results. Summary of the Invention

[0004] This invention provides a method and system for job suitability assessment based on multimodal resume parsing, the main purpose of which is to solve the problem of low accuracy in job suitability assessment based on multimodal resume parsing.

[0005] To achieve the above objectives, this invention provides a job suitability assessment method based on multimodal resume parsing, comprising: S1. Obtain the job description text and the resume file to be evaluated for the target process, and separate the resume file to be evaluated into the resume text stream and the resume image region set for the target process; S2. Use the information tuples in the resume text stream as a resume semantic archive, and convert the pixel data in the resume image region set into the image semantic credentials of the target process; S3. Based on the positional relationship between text blocks and image regions in the resume text stream on the page, associate the image semantic credentials with the corresponding text entries in the resume semantic archive to obtain the multimodal aligned resume archive of the target process; S4. Combine the skill fields and qualification fields of the job description text into a job requirement set for the target process; S5. Using the requirements in the job requirement set as an index, locate the matching fields and corresponding corroborating records in the multimodal aligned resume file to obtain the evidence set of the target process; S6. Based on the mutual corroboration strength between the hit items and the supporting text in the evidence group, generate the job suitability assessment result for the target process.

[0006] In a preferred embodiment, the step of acquiring the job description text and the resume file to be evaluated for the target process, and separating the resume file to be evaluated into a resume text stream and a resume image region set for the target process, includes: The job description in natural language and the electronic resume document of the target process are respectively used as the job description text and the resume file to be evaluated for the target process. Based on the object type identifier of the page object flow in the resume file to be evaluated, text type objects are grouped into a text object set, and image type objects are grouped into an image object set; Based on the page coordinates of the text objects in the text object set, the layout flow order between the objects is determined, and the characters corresponding to the text objects are encoded and concatenated along the layout flow order to obtain the resume text flow of the target process; The pixel sampling data and bounding box coordinates of the image object set are aggregated into a resume image region set for the target process.

[0007] In a preferred embodiment, the step of using the information tuples in the resume text stream as a resume semantic archive and converting the pixel data in the resume image region set into image semantic credentials for the target process includes: The resume text stream is segmented into words to obtain the word sequence of the target process, and the word sequence is tagged with parts of speech to obtain the part-of-speech tagging sequence of the target process. The name, time, and skill phrases of the part-of-speech tagging sequence are combined to form a resume semantic file for the target process; The visual feature vectors of the resume image region set are mapped to the image replacement text of the resume file to be evaluated in the target process to obtain the image semantic credentials of the target process.

[0008] In a preferred embodiment, the step of associating the image semantic credentials with the corresponding text entries in the resume semantic archive based on the positional relationship between text blocks and image regions in the resume text stream on the page, to obtain the multimodal aligned resume archive of the target process, includes: The page coordinates of the text blocks in the resume text stream and the page coordinates of the image regions in the resume image region set are used to determine the layout adjacency order of the text blocks and the image regions on the page. The text blocks and the image regions are then mixed and arranged according to the layout adjacency order to obtain the graphic and text layout sequence of the target process. By tracing back element by element along the starting direction of the sequence, the membership table of the target process is obtained from the image regions in the graphic layout sequence. The character content of the membership table is compared element by element with the text entries in the resume semantic file to locate the text entries corresponding to the membership table, and the image semantic credentials of the image area are attached to the text entries to obtain the image-text binding entries of the target process. The image-text binding entries and the plain text entries in the resume semantic file are merged according to the layout flow to obtain the multimodal aligned resume file of the target process.

[0009] In a preferred embodiment, the step of tracing back element by element along the sequence start direction of the image regions in the graphic layout sequence to obtain the membership table of the target process includes: Starting from the element position index value of the current image region in the graphic layout sequence, scan element by element in the decreasing direction of the element position index value to obtain the preceding text block of the target process; The image region identifier of the current image region is paired with the text block identifier of the preceding text block to obtain the membership table of the target process.

[0010] In a preferred embodiment, the step of aggregating the skill and qualification fields of the job description text into a job requirement set for the target process includes: Identify the requirement declaration blocks delimited by layout separators in the job description text to obtain the set of requirement fragments for the target process; The type identifiers in the demand fragment set are matched with the operation dimension set and the qualification dimension set respectively to obtain the skill field group and qualification field group of the target process. The skill field group and the qualification field group are then merged to obtain the job demand set of the target process.

[0011] In a preferred embodiment, the step of locating matching fields and corresponding corroborating records in the multimodal aligned resume archives using the requirements of the job demand set as an index to obtain the evidence set for the target process includes: Perform a positive maximum match between the requirement string in the job requirement set and the text fields of the image-text bound entries and plain text entries in the multimodal aligned resume file to obtain the hit entry set of the target process; The entry with the longest substring overlap in the hit entry set is marked as the index anchor entry of the target process, and the stored content of the image credential subfield in the index anchor entry is used as the verification record of the target process. Iterate the set of hit entries and the verification records according to the requirements of the job requirement set to obtain the evidence set of the target process.

[0012] In a preferred embodiment, generating the job suitability assessment result for the target process based on the mutual corroboration strength between the matching items and supporting texts in the evidence set includes: Align the hit entries and supporting text in the evidence group according to character order to obtain the character alignment array of the target process. The position of the character alignment array records the character pairs of the hit entries and the supporting text at the corresponding positions. Scan the character alignment array along the array direction, mark the same and consecutive position segments of the character pairs as homograph segments, and obtain the homograph segment list of the target process; The segment with the longest segment length in the same segment list is taken as the core segment for mutual verification, and the segment length value of the core segment for mutual verification is converted based on the total character length of the hit entry and the supporting text to obtain the mutual verification strength value of the target process. The mutual evidence strength value is attached as a strength marker to the evidence group to obtain the evidence strength ranking chain of the target process; The requirement identifier and evidence summary of the evidence strength ranking chain are structured and encapsulated to obtain the job suitability assessment result of the target process.

[0013] In a preferred embodiment, the formula for calculating the mutual verification strength value includes: in, The mutual verification strength value is... The number of characters in the matched entries. To support the evidence regarding the number of characters in the text, The number of characters in the longest consecutive overlapping character sequence between the hit entry and the supporting text. For the first The number of characters in a non-contiguous overlapping segment. This is an index for non-contiguous overlapping segments. This represents the total number of non-continuous overlapping segments. It is a natural constant.

[0014] To address the above problems, the present invention also provides a job suitability assessment system based on multimodal resume parsing, the system comprising: The data integration module acquires the job description text and the resume file to be evaluated for the target process, and separates the resume file to be evaluated into a resume text stream and a resume image region set for the target process. The image semantic credentials module uses the information tuples in the resume text stream as resume semantic archives and converts the pixel data in the resume image region set into image semantic credentials for the target process. The multimodal aligned resume archive module associates the image semantic credentials with the corresponding text entries in the resume semantic archive based on the positional relationship between text blocks and image regions in the resume text stream on the page, thereby obtaining the multimodal aligned resume archive of the target process. The job requirements module aggregates the skill and qualification fields of the job description text into a job requirements set for the target process. The evidence group module uses the requirements in the job requirement set as an index to locate the matching fields and corresponding corroborating records in the multimodal aligned resume file to obtain the evidence group for the target process. The job suitability assessment module generates the job suitability assessment result for the target process based on the mutual strength of the matching items and supporting texts in the evidence group.

[0015] Compared with the prior art, the present invention has the following beneficial effects:

[0016] 1. This technology separates the resume file to be evaluated into a resume text stream and a resume image region set, and transforms the pixel data in the image region set into image semantic credentials. Then, based on the layout position relationship, the image semantic credentials are associated with the corresponding text entries to construct a multimodal aligned resume archive. This achieves deep integration of text information and image information and precise layout-level mounting, so that visual content such as certificates and photos that were originally isolated in the resume can directly participate in the job suitability assessment as supplementary evidence of corresponding skills or experiences.

[0017] 2. This technology uses the requirements in the job demand set as an index to locate matching fields and corresponding corroborating records in multimodal alignment resume files, generates evidence groups, and calculates a quantitative job suitability assessment result based on the mutual evidence strength of the matched items and supporting texts. At the same time, the mutual evidence strength value is attached to the evidence group as a strength marker to form an evidence strength ranking chain, and finally outputs a structured and encapsulated assessment result. This realizes a complete automated assessment link from requirements to evidence to credibility ranking, which significantly improves the structuring degree and interpretability of the assessment result. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a job suitability assessment method based on multimodal resume parsing provided in an embodiment of the present invention.

[0019] Figure 2 This is a functional module diagram of a job suitability assessment system based on multimodal resume parsing provided in an embodiment of the present invention;

[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0022] This application provides a method for evaluating job suitability based on multimodal resume parsing. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0023] Reference Figure 1 The diagram shown is a flowchart illustrating a job suitability assessment method based on multimodal resume parsing according to an embodiment of the present invention. In this embodiment, the job suitability assessment method based on multimodal resume parsing includes: In this embodiment of the invention, the step of obtaining the job description text and the resume file to be evaluated for the target process, and separating the resume file to be evaluated into the resume text stream and the resume image region set for the target process, is specifically used for: The job description in natural language and the electronic resume document of the target process are respectively used as the job description text and the resume file to be evaluated for the target process. Based on the object type identifier of the page object flow in the resume file to be evaluated, text type objects are grouped into a text object set, and image type objects are grouped into an image object set; Based on the page coordinates of the text objects in the text object set, the layout flow order between the objects is determined, and the characters corresponding to the text objects are encoded and concatenated along the layout flow order to obtain the resume text flow of the target process; The pixel sampling data and bounding box coordinates of the image object set are aggregated into a resume image region set for the target process.

[0024] Specifically, the natural language paragraphs of the job description for the target process are directly used as the job description text for the target process, while the electronic resumes in portable document format provided by job seekers are directly used as the resume files to be evaluated for the target process. The natural language paragraphs of the job description are the original text content completely extracted from the job posting page or job description, and the electronic resumes are the original files without any format conversion.

[0025] Specifically, based on the page object stream stored inside the resume file to be evaluated, which is a list of objects arranged in page order, each object carries an object type identifier field, the system traverses the list and determines the type identifier value of each object.

[0026] Specifically, the page coordinates of each text object are read from the text object set. The page coordinates include the left boundary x-coordinate, top boundary y-coordinate, right boundary x-coordinate, and bottom boundary y-coordinate of the text object on the page. The system sorts all text objects according to the top boundary y-coordinate from smallest to largest. When the top boundary y-coordinates are the same, they are sorted according to the left boundary x-coordinate from smallest to largest. The sorted order is the layout flow order among the objects.

[0027] Specifically, each image object is retrieved from the image object set, and a pixel sampling operation is performed on each image object: according to the height and width of the image, the red, green and blue color component values ​​of each pixel are read row by row and column by column, and these color component values ​​are arranged into a one-dimensional array according to the reading order. This array is the pixel sampling data of the image object.

[0028] Furthermore, if the type identifier value is text, the object is added to the text object set; if the type identifier value is image, the object is added to the image object set. After traversal, two sets are obtained: the text object set and the image object set.

[0029] Furthermore, the system sequentially extracts the character sequence stored in each text object along the layout flow, concatenates these character sequences one after another in the extraction order to form a long character sequence, performs Chinese character encoding conversion on each character in the character sequence, converts each character into a corresponding binary encoding string, and then concatenates all the binary encoding strings in character order to obtain the resume text stream of the target process.

[0030] Furthermore, the bounding box coordinates are extracted from the image objects. The bounding box coordinates record the horizontal coordinate of the top left corner, the vertical coordinate of the top left corner, the horizontal coordinate of the bottom right corner, and the vertical coordinate of the bottom right corner of the image on the page. The pixel sampling data and the bounding box coordinates are stored as a set of data in a temporary list. After the above operation is performed on each image object in the image object set, all the data sets in the temporary list are gathered together to form the resume image region set of the target process.

[0031] In summary, this approach ensures the originality and integrity of the input data, avoids information loss due to format conversion or content truncation, and provides a direct and reliable source for subsequent job requirement extraction and resume parsing.

[0032] In summary, it achieves precise separation of text and images, avoids parsing interference caused by mixed storage, and allows subsequent processing of text content and image content to be carried out independently and in a targeted manner, thereby improving the accuracy and efficiency of multimodal parsing.

[0033] In summary, the reading order and logical structure of the text blocks in the original resume are fully preserved, preventing text confusion or semantic breaks caused by complex page layouts, and providing a pure text foundation with the correct order for subsequent word segmentation, part-of-speech tagging and semantic profile construction.

[0034] In summary, the original visual data and layout information of the image are preserved, which allows for the extraction of visual feature vectors at the pixel level and their association with the corresponding text regions. This provides complete and localizable image data support for multimodal alignment and the generation of image semantic credentials.

[0035] In this embodiment of the invention, the step of using the information tuples in the resume text stream as a resume semantic archive and converting the pixel data in the resume image region set into image semantic credentials for the target process is specifically used for: The resume text stream is segmented into words to obtain the word sequence of the target process, and the word sequence is tagged with parts of speech to obtain the part-of-speech tagging sequence of the target process. The name, time, and skill phrases of the part-of-speech tagging sequence are combined to form a resume semantic file for the target process; The visual feature vectors of the resume image region set are mapped to the image replacement text of the resume file to be evaluated in the target process to obtain the image semantic credentials of the target process.

[0036] Specifically, the system takes the resume text stream as a continuous character sequence as input and uses a dictionary-based maximum forward matching segmentation method. This method first loads a dictionary containing common Chinese words, professional terms, and skill keywords. Then, starting from the beginning of the character sequence, it attempts to match the longest word in the dictionary each time. If the match is successful, a word is segmented and the pointer is moved forward by the length of that word. If the match is unsuccessful, a single character is segmented as a word and the pointer is moved forward by one character. This process is repeated until the entire character sequence is completely segmented. Each word obtained after segmentation is placed into a list in the order it appears in the original text. This list is the word sequence of the target process.

[0037] Specifically, the system iterates through each word in the part-of-speech tagging sequence and its corresponding part-of-speech tag. When a word with a part-of-speech tag of a person's name or surname is encountered, the word is extracted and placed in a temporary list. When a word with a part-of-speech tag of a time word such as year, month, date, or time period is encountered, the time word is also extracted and placed in a temporary list. When a word with a part-of-speech tag of a noun is encountered and the word belongs to the predefined skill vocabulary, the skill phrase is extracted. The skill phrase may consist of multiple consecutive words. The system will check whether the current word and its subsequent words constitute a complete skill phrase.

[0038] Specifically, the system first extracts pixel sampling data for each image region from the resume image region set. For each image region, the system uses a pre-trained convolutional neural network model to extract visual feature vectors. This model consists of multiple convolutional and pooling layers stacked together. After the pixel sampling data is input into the model, it extracts visual patterns such as edges, textures, and shapes of the image through layers of convolutional operations. Finally, a fixed-length visual feature vector is output in the last fully connected layer of the model. This vector represents the visual semantic information of the image content in the form of a numerical sequence.

[0039] Furthermore, the word sequence is then tagged with parts of speech: the system adopts a sequence labeling method based on hidden Markov models, and a labeling model is pre-trained. This model can assign a part-of-speech tag such as noun, verb, adjective, person name, time word, etc. to each word according to the context features of the word. The system feeds each word in the word sequence into the model in turn. The model outputs the unique part-of-speech tag corresponding to the word based on the character composition of the current word and the part-of-speech transition probability of the words before and after it. The part-of-speech tags of all words are arranged into a tag sequence according to the word order. This sequence is the part-of-speech tagging sequence of the target process.

[0040] Furthermore, if so, the continuous multi-word combination is extracted as a whole, and then all the extracted name, time words and skill phrases in the temporary list are sorted according to their order of appearance in the original word sequence. Finally, these elements are encapsulated into a structured file, in which each entry records an information tuple. The tuple contains information types such as name, time or skill phrase and the corresponding original text content. This structured file is the resume semantic file of the target process.

[0041] Furthermore, the system then inputs this visual feature vector into a sequence-to-sequence generative model, which consists of an encoder and a decoder. The encoder compresses the visual feature vector into an intermediate semantic vector, and the decoder gradually decodes the intermediate semantic vector to generate a natural language text sequence. This text sequence is a textual description of the image content. The system uses the generated text sequence as the image replacement text for the image region. Finally, the bounding box coordinates of each image region are paired with the corresponding image replacement text and stored. All the paired data are aggregated to form the image semantic credentials of the target process.

[0042] In summary, lexical segmentation of the resume text stream yields a lexical sequence, and part-of-speech tagging of the lexical sequence yields a part-of-speech labeled sequence. This process decomposes the continuous, unstructured character stream into lexical units with independent semantics and assigns a grammatical role label to each word, providing a structured language analysis foundation for the subsequent accurate extraction of key information such as name, time, and skill phrases.

[0043] In summary, by combining names, times, and skill phrases from part-of-speech tagging sequences into a resume semantic profile, the system filters out the personnel identity information, time clues, and professional skill items most relevant to job evaluation from the original text, forming a compact and semantically clear set of core information tuples. This significantly reduces the data scale of subsequent matching calculations and improves positioning efficiency.

[0044] In summary, mapping the visual feature vectors of the resume image region set to the image replacement text of the resume file to be evaluated yields image semantic evidence. This transforms pixel data that cannot be directly used for text matching into readable natural language descriptions, enabling image content that could not originally be used for job requirement comparison to be included in the evidence chain in text form. This broadens the modal coverage of resume parsing and enhances the evidentiary value of image regions such as certificates and photos.

[0045] In this embodiment of the invention, when associating the image semantic credentials with the corresponding text entries in the resume semantic archive based on the positional relationship between text blocks and image regions in the resume text stream on the page, to obtain the multimodal aligned resume archive of the target process, it is specifically used for: The page coordinates of the text blocks in the resume text stream and the page coordinates of the image regions in the resume image region set are used to determine the layout adjacency order of the text blocks and the image regions on the page. The text blocks and the image regions are then mixed and arranged according to the layout adjacency order to obtain the graphic and text layout sequence of the target process. By tracing back element by element along the starting direction of the sequence, the membership table of the target process is obtained from the image regions in the graphic layout sequence. The character content of the membership table is compared element by element with the text entries in the resume semantic file to locate the text entries corresponding to the membership table, and the image semantic credentials of the image area are attached to the text entries to obtain the image-text binding entries of the target process. The image-text binding entries and the plain text entries in the resume semantic file are merged according to the layout flow to obtain the multimodal aligned resume file of the target process.

[0046] Specifically, based on the page coordinates carried by each text block in the resume text stream and the page coordinates carried by each image region in the resume image region set, the page coordinates include the left boundary x-coordinate, top boundary y-coordinate, right boundary x-coordinate, and bottom boundary y-coordinate of the text block or image region on the page. The system performs the first sorting of all text blocks and image regions in ascending order of the top boundary y-coordinate.

[0047] Specifically, the image regions in the graphic layout sequence are traced back element by element along the starting direction of the sequence. The system starts from the first element of the graphic layout sequence and traverses backward. Whenever an element of type image region is encountered, the position index of the image region in the current sequence is used as the backtracking starting point. The system visits the previous elements one by one in the direction of decreasing index value, i.e., the starting direction of the sequence, until the first element of type text block is encountered and the backtracking stops.

[0048] Specifically, each record in the membership table is extracted. Each record contains a text block identifier and an image region identifier. The system then searches for the corresponding text entry in the resume semantic archive based on the text block identifier.

[0049] Specifically, the system merges the image-text bound entries with the plain text entries in the resume semantic archive that are not associated with image areas in the layout flow order. The system first extracts all plain text entries from the resume semantic archive. These plain text entries are the original text entries that are not attached with image semantic credentials by any image area.

[0050] Furthermore, when the upper boundary ordinates are the same, a second sorting is performed according to the left boundary x-coordinates in ascending order. The resulting mixed sequence is the layout adjacency order within the page. The system arranges text blocks and image regions alternately in a list along this layout adjacency order. Each element in the list records whether its type is a text block or an image region and the corresponding original data. This list is the text and image layout sequence of the target process.

[0051] Furthermore, the unique identifier of the image region is paired with the unique identifier of the first preceding text block found. The pairing format is that the text block identifier points to the image region identifier. The system stores each such pairing record in a mapping table, which is the membership table of the target process.

[0052] Furthermore, the search method involves traversing each text entry in the resume semantic archive and comparing its storage location identifier with the text block identifier. When a matching text entry is found, the system binds the character content of the text entry with the current record in the membership table. At the same time, based on the image region identifier, the corresponding image replacement text is retrieved from the image semantic credentials and attached as a subfield below the found text entry, forming a combined entry containing the original text content and the additional image semantic credentials. This combined entry is the image-text binding entry for the target process.

[0053] Furthermore, the system then sorts all entries according to their layout coordinates on the original page. The sorting rule is the same as when generating the image and text layout sequence, i.e., first by the top boundary vertical coordinate and then by the left boundary horizontal coordinate. After sorting, a unified sequence is obtained, which alternately contains image and text bound entries and plain text entries. Each entry retains its original position information in the resume and the text content it carries. The semantic credentials of the attached image are also retained. This complete sequence is the multimodal aligned resume file of the target process.

[0054] In summary, the layout adjacency order is determined based on the page coordinates of text blocks in the resume text flow and the page coordinates of image regions in the resume image region set. The text blocks and image regions are then mixed and arranged according to this order to obtain the text and image layout sequence. This establishes a positional relationship between text and images that originally belong to different data structures in a unified page space order, providing a spatial basis for the correct allocation of images and text in the future.

[0055] In summary, by tracing back element by element along the starting direction of the image region in the text layout sequence, a membership table is obtained, which enables each image region to accurately find its preceding text block in the layout. This solves the problem of which text paragraph or entry title an image region belongs to in a resume, and establishes a clear pointing relationship for binding image semantic credentials with corresponding text entries.

[0056] In summary, by comparing the character content of the affiliation table with the text entries in the resume semantic archive element by element to locate the corresponding text entries, and attaching the image semantic credentials of the image area to the text entries to obtain image-text binding entries, the fusion storage of image semantic information and text semantic information in the same entry is realized. This allows image content such as certificate photos to be directly attached to the text description as supplementary evidence of the corresponding skills or experiences.

[0057] In summary, by merging the text-image bound entries with the plain text entries in the resume semantic archive according to the layout flow, a multimodal aligned resume archive is obtained. This fully preserves the layout order and content integrity of all information entries in the original resume. At the same time, each text entry that may have image corroboration is accompanied by a machine-readable image semantic description, thus constructing a structured resume representation in which text and images corroborate each other and are aligned in both position and content.

[0058] In this embodiment of the invention, when the step of tracing back element by element along the starting direction of the sequence to obtain the membership table of the target process from the image regions in the graphic layout sequence is specifically used for: Starting from the element position index value of the current image region in the graphic layout sequence, scan element by element in the decreasing direction of the element position index value to obtain the preceding text block of the target process; The image region identifier of the current image region is paired with the text block identifier of the preceding text block to obtain the membership table of the target process.

[0059] Specifically, using the element position index of the current image region in the graphic layout sequence as the backtracking starting point, the system obtains the position number of the image region in the graphic layout sequence. For example, it stores the position number of the image region in the sequence as the backtracking starting point value. Then, starting from the previous position obtained by subtracting one from the position number, it visits each element one by one in the direction of decreasing the position number, that is, towards the beginning of the sequence. Each time an element is visited, its type tag is checked.

[0060] Specifically, the image region identifier of the current image region is paired with the text block identifier of the preceding text block. The system extracts its unique identifier from the data structure of the current image region, which is called the image region identifier, and extracts its unique identifier from the data structure of the preceding text block, which is called the text block identifier. These two identifiers are combined into a key-value pair, with the text block identifier as the key and the image region identifier as the value.

[0061] Furthermore, if the type is marked as a text block, the scanning stops immediately and the text block is output as the previous neighbor text block. If the type is marked as an image region, the element is skipped and the scanning continues forward until a text block is encountered or the start position of the sequence is reached. The resulting previous neighbor text block is recorded as the previous neighbor text block of the target process.

[0062] Furthermore, this key-value pair is appended to a blank mapping table. The system repeats the above backtracking and pairing operations for each image region in the text and image layout sequence. Each time an image region is processed, a record is added to the mapping table. The complete mapping table obtained after all image regions have been processed is the membership table of the target process.

[0063] In summary, by scanning element by element from the current image region in the text layout sequence, starting from the element position index, the preceding text block is obtained in a decreasing direction. This allows each image region to uniquely identify the text block that is spatially closest to and in front of it on the page, avoiding ambiguity caused by the interleaving of images and text, and establishing a clear and traceable text reference anchor for each image region.

[0064] In summary, by pairing the image region identifier of the current image region with the text block identifier of the preceding text block to obtain a membership table, the adjacency relationship between the image and the text is transformed into a machine-readable identifier mapping record. This allows subsequent modules to directly locate the text entry to which each image region belongs by looking up the table without re-parsing the layout, significantly improving the processing efficiency and attribution accuracy of multimodal alignment.

[0065] In this embodiment of the invention, when the skill fields and qualification fields of the job description text are aggregated into the job requirement set of the target process, it is specifically used for: Identify the requirement declaration blocks delimited by layout separators in the job description text to obtain the set of requirement fragments for the target process; The type identifiers in the demand fragment set are matched with the operation dimension set and the qualification dimension set respectively to obtain the skill field group and qualification field group of the target process. The skill field group and the qualification field group are then merged to obtain the job demand set of the target process.

[0066] Specifically, the system identifies requirement declaration blocks in the job description text that are defined by formatting delimiters. Formatting delimiters include line breaks, paragraph marks, bullet points such as dots or numbers with commas, and semicolons. The system segments the job description text according to these delimiters, using line breaks as the primary delimiter. Text between two consecutive line breaks is considered an independent paragraph, and line breaks within a paragraph are considered line breaks within the same paragraph and are not segmented.

[0067] Specifically, for each requirement fragment in the requirement fragment set, the system first determines the type identifier carried by the requirement fragment. The type identifier is a tag word extracted from the requirement fragment. These tag words usually appear at the beginning of the requirement fragment or as a subheading. The system matches the extracted type identifier with two predefined reference sets. The first reference set is the operation dimension set, which contains words related to specific operation capabilities. The second reference set is the qualification dimension, which contains words related to basic admission conditions.

[0068] Furthermore, for lists using bullet points, the content following each bullet point is separately segmented into a requirement declaration block. The system traverses the entire job description text, and records the currently accumulated text fragments as a block whenever it encounters a formatting separator. Then, it clears the accumulated area and continues scanning until the end of the text. All the segmented blocks are stored sequentially into a list, which is the requirement fragment set for the target process.

[0069] Furthermore, string inclusion judgment is used during matching. If the type identifier is the same as or contains any word in the operation dimension set, the requirement fragment is assigned to the skill field group. If the type identifier is the same as or contains any word in the qualification dimension set, the requirement fragment is assigned to the qualification field group. In the case of matching two reference sets at the same time, the reference set with the higher proportion of the type identifier is assigned. After all requirement fragments have been matched, skill field group and qualification field group are obtained. The corresponding requirement fragment text is stored in the original order within each field group. Then, the skill field group and qualification field group are merged into a unified list according to the order in which they appear in the job description text. This list is the job requirement set of the target process.

[0070] In summary, identifying requirement declaration blocks delimited by layout separators in job description text yields a set of requirement fragments. This automatically breaks down unstructured natural language job descriptions into independent requirement units according to layout separators, avoiding semantic confusion caused by overlapping entire texts. This provides a clear minimum analytical granularity for the subsequent accurate classification of skill and qualification fields.

[0071] In summary, by matching the type identifiers in the demand fragment set with the operation dimension set and the qualification dimension set respectively, skill field groups and qualification field groups are obtained, and the two are merged to obtain the job demand set. This achieves automatic differentiation and classification of operation skill requirements and qualification requirements in job requirements, enabling differentiated matching strategies for different types of fields when locating resume evidence using requirements as an index. At the same time, the merged job demand set retains all the constraints in the original job description in a unified list format.

[0072] In this embodiment of the invention, when locating the matching field and corresponding corroborating record in the multimodal aligned resume file using the requirements in the job requirement set as an index to obtain the evidence set for the target process, it is specifically used for: Perform a positive maximum match between the requirement string in the job requirement set and the text fields of the image-text bound entries and plain text entries in the multimodal aligned resume file to obtain the hit entry set of the target process; The entry with the longest substring overlap in the hit entry set is marked as the index anchor entry of the target process, and the stored content of the image credential subfield in the index anchor entry is used as the verification record of the target process. Iterate the set of hit entries and the verification records according to the requirements of the job requirement set to obtain the evidence set of the target process.

[0073] Specifically, for each requirement string in the job requirement set, a positive maximum match is performed with the text fields of image-text bound entries and plain text entries in the multimodal aligned resume file. The system treats each requirement string as a pattern to be matched, and starting from the first character of the requirement string, selects a substring with the same length as the current matching window. The initial length of the matching window is set to the total number of characters in the requirement string. Then, it checks whether the substring appears completely in the text field of the current entry.

[0074] Specifically, the system compares the hit substring recorded in each entry in the hit entry set with the corresponding demand string, calculates the substring overlap length (i.e., the number of consecutive characters in the hit substring that are identical to the demand string), and selects the entry with the largest substring overlap length value from the same entries for the same demand string. The system then marks the entry with the largest substring overlap length value as the index anchor entry for the target process.

[0075] Specifically, each requirement string in the job requirement set is processed sequentially. The current requirement string is used to obtain the set of hit items to get the corresponding set of hit sub-items. Then, the index anchor item is determined from the set and the verification record is extracted. The requirement string itself, the corresponding index anchor item identifier, and the verification record are encapsulated into a data group.

[0076] Furthermore, if a match is found, the identifier of the entry and the start and end positions of the matching substring are recorded. If a match is not found, the length of the matching window is reduced by one character, and a new substring is extracted for matching. This process is repeated until the length of the matching window is reduced to a single character. After each requirement string completes the above matching for all entries, all entries that match at least one character are collected and deduplicated. These entries are grouped and stored according to the requirement string used when they were matched. Each entry records the content of the matched substring and its position in the original text, thus obtaining the set of matched entries for the target process.

[0077] Furthermore, if multiple entries have the same maximum substring overlap length, the one that appears earlier in the hit entry set is selected. Then, the image credential subfield is extracted from the data structure of the index-anchored entry. This subfield stores the content of the previously attached image semantic credential, i.e., the image replacement text. This content is extracted as is as a verification record for the requirement string. The verification record also records the original image region identifier corresponding to the image credential subfield.

[0078] Furthermore, the data set is formatted as a requirement string value plus a resume entry identifier plus an image-replaced text content. After processing one requirement string, it moves to the next requirement string, repeating the above iterative operation until all requirement strings in the job requirement set have been processed. All the obtained data sets are arranged in the original order of the requirement strings in the job requirement set and stored in a list. This list is the evidence set of the target process.

[0079] In summary, a positive maximum matching is performed on the requirement strings in the job requirement set and the text fields of image-text bound items and plain text items in the multimodal aligned resume file to obtain the set of matching items. This ensures that the longest text overlap segment can be found in the resume for each job requirement, avoiding false hits or missed hits caused by simple keyword matching, and guaranteeing the maximum coverage and accuracy of the matching results at the character level.

[0080] In summary, the entry with the longest substring overlap in the matched entry set is marked as the index anchor entry, and the stored content of the image credential subfield in that entry is used as a corroborating record. The entry most relevant to the requirement is selected from multiple possible matching results, and the semantic information of the image attached to that entry is automatically extracted as supporting evidence. This provides a reliable basis for each job requirement with the dual support of text matching and image evidence.

[0081] In summary, evidence sets are obtained by iterating through the required items and verifying records in the job demand set. Each job demand is bound to its corresponding required item and image evidence into a complete evidence unit, forming a three-level association structure from demand to resume content to supporting materials. This provides a structured and traceable evidence chain for subsequent calculation of mutual evidence strength and generation of evaluation results.

[0082] In this embodiment of the invention, when generating the job suitability assessment result for the target process based on the mutual corroboration strength between the hit items and supporting texts in the evidence group, it is specifically used for: Align the hit entries and supporting text in the evidence group according to character order to obtain the character alignment array of the target process. The position of the character alignment array records the character pairs of the hit entries and the supporting text at the corresponding positions. Scan the character alignment array along the array direction, mark the same and consecutive position segments of the character pairs as homograph segments, and obtain the homograph segment list of the target process; The segment with the longest segment length in the same segment list is taken as the core segment for mutual verification, and the segment length value of the core segment for mutual verification is converted based on the total character length of the hit entry and the supporting text to obtain the mutual verification strength value of the target process. The mutual evidence strength value is attached as a strength marker to the evidence group to obtain the evidence strength ranking chain of the target process; The requirement identifier and evidence summary of the evidence strength ranking chain are structured and encapsulated to obtain the job suitability assessment result of the target process.

[0083] Specifically, the system aligns the hit entries and supporting text in the evidence group according to character order. It extracts the complete character sequence of the hit entries and the complete character sequence of the supporting text, and aligns them position by position starting from the first character of the two sequences. At each position, it records the character pair consisting of two characters.

[0084] Specifically, scan each position in the character alignment array along the array direction, i.e. from left to right, and check whether the two characters in the character pair stored at that position are the same and both characters are not empty. If they are the same, start marking a segment with the same character and continue scanning forward. As long as the character pairs at subsequent positions continuously satisfy the condition that the two characters are the same and both are not empty, continue to expand the current segment with the same character.

[0085] Specifically, the system compares the segment length values ​​of all identical segments in the identical segment list and selects the segment with the largest segment length value as the core segment for mutual verification. Then, the system calculates the total number of characters in the hit entries and the total number of characters in the supporting text. The system divides the square root of the product of the segment length value of the core segment for mutual verification with the total number of characters in the hit entries and the total number of characters in the supporting text to obtain an intermediate ratio. At the same time, the system removes all identical segments from the identical segment list except for the core segment for mutual verification.

[0086] Specifically, the mutual verification strength value is appended as a strength marker to each evidence entry in the evidence group. Each evidence entry originally contains a requirement string, an index anchor entry identifier, and a corroboration record. The system adds the calculated mutual verification strength value in numerical form to the end of the entry as a new field.

[0087] Specifically, the system extracts the requirement identifier (i.e., the content of the requirement string) and the evidence summary (i.e., a brief description of the index anchor entry identifier and the corroborating record) from each evidence item in the evidence strength ranking chain. The system then organizes this information into a structured data object according to the order of the ranking chain.

[0088] Furthermore, if the hit entry is longer than the supporting text, the excess positions of the supporting text are filled with empty characters; if the supporting text is longer, the excess positions of the hit entry are filled with empty characters. All character pairs are arranged in order into a two-dimensional array. The number of rows in this array is equal to the number of columns, which is equal to the number of characters in the longer sequence. Each position in the array stores a character pair. This array is the character alignment array of the target process.

[0089] Furthermore, once a character pair is encountered that is not identical or any character is empty, the current segment with the same character is terminated and the start and end positions of the segment, as well as the segment length (i.e., the number of consecutive identical characters), are recorded. Then, the subsequent positions are scanned to find the next segment with the same character. After scanning the entire array, a list of records of all segments with the same character is obtained. This list is the segment list of the target process.

[0090] Furthermore, the segment length values ​​of these identical segments are summed to obtain the total number of characters in the non-contiguous overlapping segments. This total number of characters is divided by the sum of the total number of characters in the hit entries and the total number of characters in the supporting text to obtain an additional ratio. This additional ratio is added to the previous intermediate ratio to obtain a product value. The ratio of the natural constant to the sum of the total number of characters in the hit entries and the total number of characters in the supporting text is then calculated. This ratio is added to the previous ratio to obtain a denominator. The product value is divided by this denominator to obtain the final value, which is the mutual evidence strength value of the target process.

[0091] Furthermore, all evidence items are then sorted in descending order of mutual evidence strength value. When the mutual evidence strength values ​​are the same, they are arranged according to the original order of the requirement strings in the job requirement set. After sorting, a new list of evidence items is obtained, which is the evidence strength sorting chain of the target process.

[0092] Furthermore, the object contains a root node and several child nodes. Each child node corresponds to an evidence entry and stores its requirement identifier, mutual evidence strength value, and evidence summary. Finally, this structured data object is serialized into a transmittable text format. The serialized object is the job suitability assessment result of the target process.

[0093] In summary, aligning the hit entries in the evidence group with the supporting text in character order to obtain a character-aligned array establishes a precise correspondence between the two text sequences at each character position, providing a position-by-position comparable two-dimensional data structure for subsequent identification of identical character segments and calculation of overlap length.

[0094] In summary, scanning the character alignment array along the array direction and marking the same and consecutive position segments of character pairs as identical segments yields a list of identical segments. All consecutive identical character segments are automatically extracted from the character-by-character alignment results, making the identification process of the longest consecutive overlapping sequence and non-consecutive overlapping segments quantifiable and reproducible.

[0095] In summary, the segment with the longest segment length in the same segment list is taken as the core segment for mutual verification. The segment length value is converted based on the total character length of the hit entries and supporting texts to obtain the mutual verification strength value. The relationship between the longest continuous overlapping length and the overall length of both texts, as well as the contribution of all non-continuous overlapping segments, are quantified into a unified value, making the degree of verification between texts of different lengths comparable.

[0096] In summary, by attaching the mutual evidence strength value as a strength marker to the evidence group, an evidence strength ranking chain is obtained. This assigns a ranking confidence measure to the evidence items corresponding to each job requirement, enabling the evaluation system to organize evidence from high to low reliability, which facilitates the subsequent presentation of the most reliable matching results.

[0097] In summary, the structured encapsulation of the requirement identifiers and evidence summaries in the evidence strength ranking chain yields the job fit assessment results. The original hit items, supporting texts, image credentials, and mutual evidence strength values ​​are integrated into concise and clear requirement evidence pairs, which are output in a standardized format for use by recruiters or downstream systems. This achieves a complete closed loop from multimodal resume parsing to quantifiable job fit assessment.

[0098] In this embodiment of the invention, the formula for calculating the mutual verification strength value is specifically used for: in, The mutual verification strength value is... The number of characters in the matched entries. To support the evidence regarding the number of characters in the text, The number of characters in the longest consecutive overlapping character sequence between the hit entry and the supporting text. For the first The number of characters in a non-contiguous overlapping segment. This is an index for non-contiguous overlapping segments. This represents the total number of non-continuous overlapping segments. It is a natural constant.

[0099] Specifically, the number of characters in the hit entries comes from the length statistics of the complete character sequences of the hit entries in the evidence group; the number of characters in the supporting text comes from the length statistics of the complete character sequences of the supporting text in the evidence group; the number of characters in the longest consecutive overlapping character sequence comes from the segment length value of the core mutual evidence segment with the largest segment length in the same segment list; the number of characters in the non-consecutive overlapping segments comes from the segment length value of each other same segment in the same segment list except for the core mutual evidence segment; the total number of non-consecutive overlapping segments comes from the number of other same segments in the same segment list except for the core mutual evidence segment; and the natural constant is a fixed transcendental number in mathematics.

[0100] Furthermore, the mutual verification strength value calculated by the formula is used to measure the reliability of the corroboration between the hit item and the supporting text. The higher the mutual verification strength value, the higher the degree of overlap between the hit item and the supporting text at the character level and the denser the distribution of overlapping segments. The formula reflects the core consistency by the ratio of the longest continuous overlapping length to the square root of the total length of both parties, reflects the supplementary contribution of other sporadic consistent information by the ratio of the total length of non-continuous overlapping segments to the total length of both parties, and penalizes excessively short texts by using the reciprocal factor formed by the ratio of the natural constant to the total length of both parties to prevent false overestimation.

[0101] In general, the mutual evidence strength value reaches its theoretical maximum when the matched item and the supporting text are exactly the same and have no difference. The mutual evidence strength value increases when the longest continuous overlapping length or the total length of non-continuous overlapping segments increases. The mutual evidence strength value decreases when the total length of both texts increases while the overlapping length remains unchanged. When the total length of both texts is very small, the influence of the natural constant causes the reciprocal factor in the denominator to significantly suppress the mutual evidence strength value.

[0102] Compared with the prior art, the present invention has the following beneficial effects:

[0103] 1. This technology separates the resume file to be evaluated into a resume text stream and a resume image region set, and transforms the pixel data in the image region set into image semantic credentials. Then, based on the layout position relationship, the image semantic credentials are associated with the corresponding text entries to construct a multimodal aligned resume archive. This achieves deep integration of text information and image information and precise layout-level mounting, so that visual content such as certificates and photos that were originally isolated in the resume can directly participate in the job suitability assessment as supplementary evidence of corresponding skills or experiences.

[0104] 2. This technology uses the requirements in the job demand set as an index to locate matching fields and corresponding corroborating records in multimodal alignment resume files, generates evidence groups, and calculates a quantitative job suitability assessment result based on the mutual evidence strength of the matched items and supporting texts. At the same time, the mutual evidence strength value is attached to the evidence group as a strength marker to form an evidence strength ranking chain, and finally outputs a structured and encapsulated assessment result. This realizes a complete automated assessment link from requirements to evidence to credibility ranking, which significantly improves the structuring degree and interpretability of the assessment result.

[0105] like Figure 2 The diagram shown is a functional module diagram of a job suitability assessment system based on multimodal resume parsing provided in an embodiment of the present invention.

[0106] The job suitability assessment system 100 based on multimodal resume parsing described in this invention can be installed in an electronic device. Depending on the functions implemented, the job suitability assessment system 100 may include a data integration module 101, an image semantic credential module 102, a multimodal aligned resume archive module 103, a job requirement module 104, an evidence group module 105, and a job suitability assessment module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.

[0107] In this embodiment, the functions of each module / unit are as follows: The data integration module acquires the job description text and the resume file to be evaluated for the target process, and separates the resume file to be evaluated into a resume text stream and a resume image region set for the target process. The image semantic credentials module uses the information tuples in the resume text stream as resume semantic archives and converts the pixel data in the resume image region set into image semantic credentials for the target process. The multimodal aligned resume archive module associates the image semantic credentials with the corresponding text entries in the resume semantic archive based on the positional relationship between text blocks and image regions in the resume text stream on the page, thereby obtaining the multimodal aligned resume archive of the target process. The job requirements module aggregates the skill and qualification fields of the job description text into a job requirements set for the target process. The evidence group module uses the requirements in the job requirement set as an index to locate the matching fields and corresponding corroborating records in the multimodal aligned resume file to obtain the evidence group for the target process. The job suitability assessment module generates the job suitability assessment result for the target process based on the mutual strength of the matching items and supporting texts in the evidence group.

[0108] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0109] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0110] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0111] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0112] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for evaluating job suitability based on multimodal resume parsing, characterized in that, The method includes: S1. Obtain the job description text and the resume file to be evaluated for the target process, and separate the resume file to be evaluated into the resume text stream and the resume image region set for the target process; S2. Use the information tuples in the resume text stream as a resume semantic archive, and convert the pixel data in the resume image region set into the image semantic credentials of the target process; S3. Based on the positional relationship between text blocks and image regions in the resume text stream on the page, associate the image semantic credentials with the corresponding text entries in the resume semantic archive to obtain the multimodal aligned resume archive of the target process; S4. Combine the skill fields and qualification fields of the job description text into a job requirement set for the target process; S5. Using the requirements in the job requirement set as an index, locate the matching fields and corresponding corroborating records in the multimodal aligned resume file to obtain the evidence set of the target process; S6. Based on the mutual corroboration strength between the hit items and the supporting text in the evidence group, generate the job suitability assessment result for the target process.

2. The job suitability assessment method based on multimodal resume parsing as described in claim 1, characterized in that, The process of acquiring the job description text and the resume file to be evaluated for the target process, and separating the resume file to be evaluated into the resume text stream and the resume image region set for the target process, includes: The job description in natural language and the electronic resume document of the target process are respectively used as the job description text and the resume file to be evaluated for the target process. Based on the object type identifier of the page object flow in the resume file to be evaluated, text type objects are grouped into a text object set, and image type objects are grouped into an image object set; Based on the page coordinates of the text objects in the text object set, the layout flow order between the objects is determined, and the characters corresponding to the text objects are encoded and concatenated along the layout flow order to obtain the resume text flow of the target process; The pixel sampling data and bounding box coordinates of the image object set are aggregated into a resume image region set for the target process.

3. The job suitability assessment method based on multimodal resume parsing as described in claim 1, characterized in that, The step of using information tuples from the resume text stream as a resume semantic archive and converting pixel data from the resume image region set into image semantic credentials for the target process includes: The resume text stream is segmented into words to obtain the word sequence of the target process, and the word sequence is tagged with parts of speech to obtain the part-of-speech tagging sequence of the target process. The name, time, and skill phrases of the part-of-speech tagging sequence are combined to form a resume semantic file for the target process; The visual feature vectors of the resume image region set are mapped to the image replacement text of the resume file to be evaluated in the target process to obtain the image semantic credentials of the target process.

4. The job suitability assessment method based on multimodal resume parsing as described in claim 1, characterized in that, The step of associating the image semantic credentials with the corresponding text entries in the resume semantic archive based on the positional relationship between text blocks and image regions in the resume text stream on the page, to obtain the multimodal aligned resume archive of the target process, includes: The page coordinates of the text blocks in the resume text stream and the page coordinates of the image regions in the resume image region set are used to determine the layout adjacency order of the text blocks and the image regions on the page. The text blocks and the image regions are then mixed and arranged according to the layout adjacency order to obtain the graphic and text layout sequence of the target process. By tracing back element by element along the starting direction of the sequence, the membership table of the target process is obtained from the image regions in the graphic layout sequence. The character content of the membership table is compared element by element with the text entries in the resume semantic file to locate the text entries corresponding to the membership table, and the image semantic credentials of the image area are attached to the text entries to obtain the image-text binding entries of the target process. The image-text binding entries and the plain text entries in the resume semantic file are merged according to the layout flow to obtain the multimodal aligned resume file of the target process.

5. The job suitability assessment method based on multimodal resume parsing as described in claim 4, characterized in that, The step of tracing back element by element along the starting direction of the sequence through the image regions in the graphic layout sequence to obtain the membership table of the target process includes: Starting from the element position index value of the current image region in the graphic layout sequence, scan element by element in the decreasing direction of the element position index value to obtain the preceding text block of the target process; The image region identifier of the current image region is paired with the text block identifier of the preceding text block to obtain the membership table of the target process.

6. The job suitability assessment method based on multimodal resume parsing as described in claim 1, characterized in that, The step of compiling the skill and qualification fields of the job description text into the job requirement set for the target process includes: Identify the requirement declaration blocks delimited by layout separators in the job description text to obtain the set of requirement fragments for the target process; The type identifiers in the demand fragment set are matched with the operation dimension set and the qualification dimension set respectively to obtain the skill field group and qualification field group of the target process. The skill field group and the qualification field group are then merged to obtain the job demand set of the target process.

7. The job suitability assessment method based on multimodal resume parsing as described in claim 1, characterized in that, The process of locating matching fields and corresponding corroborating records in the multimodal aligned resume archives using the requirements in the job demand set as an index, thereby obtaining the evidence set for the target process, includes: Perform a positive maximum match between the requirement string in the job requirement set and the text fields of the image-text bound entries and plain text entries in the multimodal aligned resume file to obtain the hit entry set of the target process; The entry with the longest substring overlap in the hit entry set is marked as the index anchor entry of the target process, and the stored content of the image credential subfield in the index anchor entry is used as the verification record of the target process. Iterate the set of hit entries and the verification records according to the requirements of the job requirement set to obtain the evidence set of the target process.

8. The job suitability assessment method based on multimodal resume parsing as described in claim 1, characterized in that, The step of generating the job suitability assessment result for the target process based on the mutual corroboration strength between the matching items and supporting texts in the evidence group includes: Align the hit entries and supporting text in the evidence group according to character order to obtain the character alignment array of the target process. The position of the character alignment array records the character pairs of the hit entries and the supporting text at the corresponding positions. Scan the character alignment array along the array direction, mark the same and consecutive position segments of the character pairs as homograph segments, and obtain the homograph segment list of the target process; The segment with the longest segment length in the same segment list is taken as the core segment for mutual verification, and the segment length value of the core segment for mutual verification is converted based on the total character length of the hit entry and the supporting text to obtain the mutual verification strength value of the target process. The mutual evidence strength value is attached as a strength marker to the evidence group to obtain the evidence strength ranking chain of the target process; The requirement identifier and evidence summary of the evidence strength ranking chain are structured and encapsulated to obtain the job suitability assessment result of the target process.

9. The job suitability assessment method based on multimodal resume parsing as described in claim 8, characterized in that, The formula for calculating the mutual verification strength value includes: in, The mutual verification strength value is... The number of characters in the matched entries. To support the evidence regarding the number of characters in the text, The number of characters in the longest consecutive overlapping character sequence between the hit entry and the supporting text. For the first The number of characters in a non-contiguous overlapping segment. This is an index for non-contiguous overlapping segments. This represents the total number of non-continuous overlapping segments. It is a natural constant.

10. A job suitability assessment system based on multimodal resume parsing, used to implement the job suitability assessment method based on multimodal resume parsing as described in any one of claims 1-9, characterized in that, The system includes: The data integration module acquires the job description text and the resume file to be evaluated for the target process, and separates the resume file to be evaluated into a resume text stream and a resume image region set for the target process. The image semantic credentials module uses the information tuples in the resume text stream as resume semantic archives and converts the pixel data in the resume image region set into image semantic credentials for the target process. The multimodal aligned resume archive module associates the image semantic credentials with the corresponding text entries in the resume semantic archive based on the positional relationship between text blocks and image regions in the resume text stream on the page, thereby obtaining the multimodal aligned resume archive of the target process. The job requirements module aggregates the skill and qualification fields of the job description text into a job requirements set for the target process. The evidence group module uses the requirements in the job requirement set as an index to locate the matching fields and corresponding corroborating records in the multimodal aligned resume file to obtain the evidence group for the target process. The job suitability assessment module generates the job suitability assessment result for the target process based on the mutual strength of the matching items and supporting texts in the evidence group.