Method and device for positioning signature area in contract document and electronic equipment
By combining text analysis and image processing methods, the machine learning model is used to calculate semantic correlation and feature similarity, and the signature area in the electronic contract is determined, which solves the problem of positioning difficulties in traditional methods and achieves efficient and accurate automated positioning.
Patent Information
- Application Number
- CN202510156782.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-20
AI Technical Summary
It is difficult to locate the signature area in electronic contracts, especially when the contract is long or the signature style is complex, it is difficult for traditional methods to identify the signature area efficiently and accurately.
Using a method combining text analysis and image processing, the semantic correlation between text units and preset keywords is calculated through the first machine learning model, and the target text area is determined; then, the second machine learning model is used to extract image features, calculate feature similarity, determine the candidate layout unit area, and finally determine the signature area based on the position overlap.
It improves the accuracy and efficiency of the positioning of the signature area, realizes the automatic positioning of the signature area, and reduces the dependence on manual search.
Smart Images

Figure CN120182977A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of text processing, and particularly to a method and device for positioning signature areas in contract documents and an electronic device. Background Art
[0002] With the gradual transformation of contract management towards electronicization, the storage and processing efficiency of contract documents has been significantly improved. However, electronic contracts usually exist in unstructured formats such as PDF and scanned images. The positions and forms of signature areas are diverse and may be scattered on different pages of the document, which poses new challenges for quickly locating signature areas. Traditional keyword search or manual page-turning methods are difficult to efficiently and accurately identify signature areas, especially in the case of long contract documents or complex signature styles, where omissions or misjudgments are likely to occur. In addition, signatures in electronic contracts may exist in various forms such as handwritten signatures, electronic signatures, or seal images, further increasing the difficulty of automatic recognition.
[0003] Therefore, a solution that can accurately locate signature areas in contract documents with a high accuracy rate is needed. Summary of the Invention
[0004] Embodiments of this application provide a method and device for positioning signature areas in contract documents and an electronic device to solve the defect of low efficiency relying on manual search in the prior art.
[0005] To achieve the above technical objectives, embodiments of this application propose a method for positioning signature areas in contract documents, including:
[0006] Obtain a target contract document;
[0007] Use a first machine learning model to calculate a first semantic relevance between each text unit in the target contract document and a preset keyword, where the preset keyword is a word related to signing and / or stamping;
[0008] Determine a target text unit whose calculated first semantic relevance is greater than a preset semantic threshold;
[0009] Determine a target text area corresponding to each target text unit according to the text unit position information of the target text unit, where the target text area is a rectangular area, and the rectangular area has the text unit position as a starting vertex, and extends a first distance and a second distance from the text unit position towards a first direction and a second direction perpendicular to the first direction respectively as two adjacent sides of the rectangular area, thereby defining the rectangular area;
[0010] Use a second machine learning model to extract multiple image features of the target contract document;
[0011] Generate position features of each of the multiple image features relative to other image features based on the positional relationships among the multiple image features;
[0012] Calculate the feature similarity between each position feature and each preset layout feature;
[0013] Determine, as candidate layout unit regions, the image features corresponding to the position features for which the calculated feature similarity is greater than the preset layout threshold corresponding to each preset layout feature, where the candidate layout unit regions are the preset layout unit regions corresponding to the preset layout features;
[0014] Calculate the second semantic relevance between the classification information of the preset layout unit regions corresponding to each candidate layout unit region and the preset keyword;
[0015] Determine, as target layout unit regions, the candidate layout unit regions for which the calculated second semantic relevance is greater than the preset semantic threshold;
[0016] Calculate the position overlap degree between the target text region and the target layout unit region;
[0017] Determine, as signature and seal regions, the target text regions for which the calculated position overlap degree is greater than the preset overlap threshold;
[0018] Another embodiment of the present application proposes a device for positioning signature and seal regions in a contract document, including:
[0019] An acquisition module, configured to acquire a target contract document;
[0020] A text parsing module, configured to use a first machine learning model to calculate the first semantic relevance between each text unit in the target contract document and a preset keyword, where the preset keyword is a word related to signing and / or stamping; determine, as target text units, the text units for which the calculated first semantic relevance is greater than the preset semantic threshold; and determine, according to the text unit position information of the target text units, target text regions corresponding to the target text units, where the target text regions are rectangular regions, where the rectangular regions have the text unit positions as starting vertices, and extend a first distance and a second distance from the text unit positions in a first direction and a second direction perpendicular to the first direction, respectively, as two adjacent sides of the rectangular regions, thereby defining the rectangular regions;
[0021] An image processing module, configured to use a second machine learning model to extract multiple image features of the target contract document; generate position features of each of the multiple image features relative to other image features based on the positional relationship between the multiple image features; calculate the feature similarity between each position feature and each preset layout feature; determine, as a candidate layout unit area, an image feature corresponding to a position feature for which the calculated feature similarity is greater than a preset layout threshold corresponding to each preset layout feature, where the candidate layout unit area is a preset layout unit area corresponding to the preset layout feature; calculate a second semantic relevance between the classification information of the preset layout unit area corresponding to each candidate layout unit area and the preset keyword; and determine, as a target layout unit area, a candidate layout unit area for which the calculated second semantic relevance is greater than a preset semantic threshold.
[0022] A determination module, configured to calculate a position overlap degree between the target text area and the target layout unit area; and determine, as a signature area, a target text area for which the calculated position overlap degree is greater than a preset overlap threshold.
[0023] An embodiment of the present application further provides an electronic device, including:
[0024] A memory, configured to store a program;
[0025] A processor, configured to run the program stored in the memory to execute a method for positioning a signature area in a contract document according to an embodiment of the present application.
[0026] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program executable by a processor is stored, where the program, when executed by the processor, implements a method for positioning a signature area in a contract document as provided in an embodiment of the present application.
[0027] According to the method, device, and electronic device for positioning a signature area in a contract document according to an embodiment of the present application, by performing text parsing on a target contract document to determine a target text area related to a preset keyword, then using a second machine learning model to extract image features of the target contract document, and calculating position features of the image features relative to other image features, candidate layout unit areas are determined based on the feature similarity with preset layout features, and then target layout units are determined based on the semantic relevance between the candidate layout units and the preset keyword, and finally, signature areas are determined according to the position overlap degree between the target text area and the target layout unit. Therefore, the signature area positioning solution according to the embodiment of the present application greatly improves the accuracy of signature area positioning by combining the target area determined by text parsing and the target area determined by image processing, enables automatic positioning of signature areas, and improves the positioning efficiency of signature areas.
[0028] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specific embodiments of this application are specifically given. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0030] Figure 1 is a flowchart of an embodiment of the signature area positioning method in the contract document provided by this application;
[0031] Figure 2 is a schematic structural diagram of an embodiment of the signature area positioning device in the contract document provided by this application;
[0032] Figure 3 is a schematic structural diagram of an embodiment of the electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The exemplary embodiments of this application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that this application can be more thoroughly understood and the scope of this application can be fully communicated to those skilled in the art.
[0034] Embodiment 1
[0035] As contract management moves fully towards electronicization, the digital storage and circulation of contract documents have become the mainstream trend. However, electronic contracts are usually saved in unstructured formats such as PDF and images, and the signature areas therein often lack unified format and position standards, making it extremely difficult to quickly locate the signature areas. Traditional keyword search methods are difficult to effectively identify image-based signatures or seals, and the manual search method is inefficient and error-prone when faced with a large number of electronic contracts. In addition, the signature forms in electronic contracts are diverse, including handwritten signatures, electronic signatures, seal images, etc., and may be distributed on different pages or positions of the document, further increasing the complexity of automatic recognition.
[0036] For this reason, the embodiments of this application provide a signature area positioning solution in the contract document, as Figure 1 shown in Figure 1The flowchart of an embodiment of the signature area positioning method in the contract document provided by this application. As Figure 1 shown, in the signature area positioning method in the contract document of the embodiment of this application, it may include:
[0037] S101: Obtain the target contract document.
[0038] According to the signature area positioning method of the contract document provided by the embodiment of this application, in step S101, the server may obtain the specified target contract document. The target contract documents in the embodiments of this application may cover various types, such as commercial contracts, labor contracts, lease contracts, sales contracts, etc. It should be noted that with the increasing complexity of contract services, a contract may involve multiple parties, such as Party A, Party B, Party C, etc. in a commercial contract. In view of this situation, in the signature area positioning method of the embodiments of this application, the historical contract processing data may be read and the relevant contract documents may be parsed, or the contract documents pre-stored in a separate server or a server deploying the method of the embodiments of this application may be obtained. In this way, the target contract document can be efficiently parsed or directly obtained in step S101, providing a basis for subsequent signature area positioning.
[0039] S102: Use the first machine learning model to calculate the first semantic relevance between each text unit in the target contract document and a preset keyword.
[0040] In step S102, the first machine learning model may be used to parse the target contract document obtained in step S101. For example, in the embodiments of the application, a semantic-based machine learning model may be used to parse the target contract document to obtain the word vector of each word, and then a word vector model may be used to calculate the semantic similarity between each word vector and the preset keyword. In particular, in the embodiments of this application, the preset keyword may be a word related to signing and / or stamping.
[0041] S103: Determine the text units with the calculated first semantic relevance greater than the preset semantic threshold as target text units.
[0042] In step S103, the first semantic relevance degrees between each text unit calculated in step S102 and a preset keyword can be sorted, and the text units with semantic relevance degrees greater than a preset semantic threshold are determined as target text units. For example, in the embodiment of the present application, all text units can be sorted in descending order according to the semantic relevance degrees calculated in step S102, so that the text units greater than the preset semantic threshold are regarded as the text units with the closest semantic association with the preset keyword, and thus are determined as target text units. In the embodiment of the present application, these target text units may include key information related to contract signature and seal, such as words containing "signature", "seal", "signing party", etc. or words semantically similar thereto.
[0043] S104: Determine a target text area corresponding to each target text unit according to the text unit position information of the target text unit.
[0044] In step S104, the position information of the target text unit determined in step S103 can be obtained. For example, the coordinate information of the first text unit or the last text unit of the target text unit can be obtained as its position information, so that the area range where a signature or seal may exist can be further determined based on the target text unit. For example, in the embodiment of the present application, the target text unit may be a text unit containing the words "signature" or "seal", and therefore, the area where the signature or seal finally exists is likely to be near this text unit. For example, usually, the user will sign or seal in the area on the right side or below the area where the words "signature" or "seal" are marked in the contract. Therefore, when the target text unit containing the words "signature" or "seal" is determined in step S103, in step S104, the target text area can be formed by extending based on the coordinates of the first word or the last word in such a target text unit determined in step S103. For example. The coordinates of the first word or the last word of the target text unit can be used as the starting point, and for example, the first distance and the second distance can be extended in the first direction and the second direction respectively to form two adjacent sides of a rectangle, and correspondingly, the first distance and the second distance are continued to form the other two sides of the rectangle, so as to form a rectangular area with the target text unit as the vertex as the target text area. For example, usually, the writing direction of text is from left to right, and the arrangement direction of text is from top to bottom. Therefore, the first direction can be the right side direction on the page of the target contract document, and the second direction can be the lower side direction on the page of the target contract document, so that the lower right side of the target text unit can be used as the target text area. Of course, depending on the different writing texts, the text direction of the contract document can also be different, and in the embodiment of the present application, the above text direction can be determined by parsing the target text unit of the contract document. For example, according to the semantic vector of the target text unit calculated by the first machine learning model, the semantic arrangement order of the words contained in the target text unit can be determined. For example, when the target text unit contains the two words "signature", the semantic arrangement order of the two words "sign" and "ature" contained therein can be determined according to the semantic vector, that is, the line direction of the words is from left to right, as the first direction. Then, the second machine learning model can be used to calculate the semantic logical relationship between the target text unit and other line text units or other paragraph text units to determine the second direction, that is, based on the semantic logical relationship between the target text unit and the text units in the adjacent line or the text units in the adjacent paragraph, to determine the arrangement relationship between different lines / paragraphs in the target contract document, so as to determine the second direction.Then, in step S104, the value of the first distance can also be determined according to the position information of adjacent characters in the target text unit. For example, the character distance can be calculated based on the coordinates of adjacent characters, that is, the length of each character, and then the length of each character is multiplied by a first multiple, such as 10 characters, and the resulting value is used as the value of the first distance. The line distance between the target text unit and other line text units and the paragraph distance between the target text unit and other paragraph text units are calculated according to the position information of the target text unit, the line text unit, and the paragraph text unit, and the maximum value of the line distance and the paragraph distance is used as the value of the second distance. Therefore, a rectangular area can be determined with the target text unit as the starting vertex and extending the first distance in the direction of character arrangement, that is, the line direction, and extending the second distance in the direction of paragraph arrangement. In other words, in the embodiments of the present application, the position of the text unit can be used as the starting vertex of the rectangular area, and the first distance and the second distance are extended from the position of the text unit in the first direction and the second direction perpendicular to the first direction respectively as the two adjacent sides of the rectangular area, thereby defining the rectangular area.
[0045] In addition, in the embodiments of the present application, the shape of the target text area is not limited to a rectangular shape, and other shapes can also be used. For example, the target text unit can be used as the origin, and a circular area can be defined with the first distance or the second distance as the radius as the target text area.
[0046] S105: Use a second machine learning model to extract multiple image features of the target contract document.
[0047] S106: Generate the position features of each image feature in the multiple image features relative to other image features based on the position relationship between the multiple image features.
[0048] In step S105, for example, a convolutional neural network model can be used to extract multiple image features from the target contract document. Then, in step S106, a convolutional neural network can also be used to determine the position relationship between the multiple image features extracted in step S105, thereby generating the position features of each image feature relative to other image features.
[0049] Before extracting the image features in step S105, the target contract document can also be preprocessed first. For example, grayscale processing can be performed first to convert the color image into a grayscale image, which can reduce the data processing volume in subsequent steps. In addition, a filtering algorithm, such as Gaussian filtering, can be used to remove the noise in the converted grayscale image to make the image clearer and facilitate subsequent processing. In addition, the contrast of the converted grayscale image can be further enhanced through image enhancement techniques, such as histogram equalization, to highlight the features of signatures and seals.
[0050] S107: Calculate the feature similarity between each position feature and each preset layout feature.
[0051] S108: Determine the image features corresponding to the position features whose calculated feature similarity is greater than the preset layout threshold corresponding to each preset layout feature as candidate layout unit regions.
[0052] In step S107, the feature similarity between the position features generated in step S106 and the preset layout features can be calculated. Then in step S108, the candidate layout units can be further determined based on the feature similarity calculated in step S107. For example, in step S108, the similarities can be sorted first, and the position information of the image features greater than the preset threshold can be further obtained, and the region containing a predetermined number or more of such image features greater than the preset threshold can be determined as the candidate layout unit region. In the embodiments of the present application, the candidate layout unit region is the preset layout unit region corresponding to a certain preset layout feature. For example, in the embodiments of the present application, preset layout unit regions such as a title region, a body region, and a signature region can be preset. In step S107, the position features of each image feature can be calculated for similarity with these preset layout unit regions to determine which layout unit the image feature belongs to, and then the image features with similarities greater than the preset threshold can be summarized, and the coordinates of the outermost image features among these image features can be selected as the boundaries of the candidate layout unit region based on the position information of these image features to generate the candidate layout unit region. In addition, in the embodiments of the present application, according to the region information of the preset layout unit regions corresponding to these image features, such as region size information, the center of these image features can be used as the region center, and the region size information of the corresponding preset layout unit region can be used as the size of the candidate layout unit region to generate the candidate layout unit region.
[0053] S109: Calculate the second semantic relevance between the classification information of the preset layout unit region corresponding to each candidate layout unit and the preset keywords.
[0054] In step S109, for each image feature determined in step S108 to belong to a preset layout unit, its classification information can be obtained, for example, belonging to one of the classifications such as a title area, a body text area, a signature area, etc. Then, the semantic relevance between the obtained classification information and the preset keyword used in step S102 can be calculated, that is, whether the classification corresponding to the image feature determined in step S108 is closely related to the preset keyword. For example, if the image feature determined in step S108 belongs to the title area, and the preset keyword used in step S102 is "signature", then in step S109, the semantic relevance between the classification "title" to which the image feature belongs and "signature" can be calculated.
[0055] S110: Determine the candidate layout unit area with the calculated second semantic relevance greater than the preset semantic threshold as the target layout unit area.
[0056] S111: Calculate the position overlap degree between the target text area and the target layout unit area.
[0057] S112: Determine the target text area with the calculated position overlap degree greater than the preset overlap threshold as the signature area.
[0058] In step S110, ranking can be performed based on the semantic relevance between the classification information to which each image feature belongs calculated in step S109 and the preset keyword, and the candidate layout unit area with a semantic relevance greater than the preset semantic threshold is determined as the target layout unit area. Then, in step S111, the position overlap degree between the target layout unit area determined in step S110 and the target text area determined in step S104 can be calculated. For example, the position overlap degree can be calculated based on the boundary coordinate information of the target layout unit area and the boundary coordinate information of the target text area, and in step S112, the target text area with an overlap degree greater than the preset threshold can be determined as the signature area. That is, the target text area determined based on text parsing and the target layout unit area determined based on image processing are combined, and when the overlap degree of the areas determined by the two methods is relatively high, it can be confirmed that the possibility of this area belonging to the signature area is also relatively high. Therefore, it is possible to avoid misrecognition caused by determining the signature area by text parsing or image processing methods, thereby improving the accuracy of signature area recognition.
[0059] In addition, in the embodiments of the present application, when determining the target layout unit area based on the candidate layout unit areas where the calculated second semantic relevance is greater than the preset semantic threshold, the image features included therein can be extracted from the candidate layout unit areas where the calculated second semantic relevance is greater than the preset semantic threshold by using a convolutional neural network, and then the feature matching degree between the extracted image features and the preset keyword features is calculated; and the candidate layout unit containing the image features with a feature matching degree greater than the preset matching threshold is determined as the target layout unit.
[0060] In addition, in the embodiments of the present application, after the signature area is determined in step S112, the signature area positioning method according to the embodiments of the present application may further include: extracting the image features included therein from the signature area determined in step S112 by using a convolutional neural network; calculating the feature matching degree between the extracted image features and the preset keyword features; determining the candidate key features with a feature matching degree greater than the preset matching threshold; calculating the position information of the edge pixels of the candidate key features; and updating the boundary of the signature area determined in step S112 according to the position information.
[0061] According to the signature area positioning method in the contract document of the embodiments of the present application, the target text area related to the preset keyword is determined by parsing the text of the target contract document, and then the image features of the target contract document are extracted by using the second machine learning model, and the position features of the image features relative to other image features are calculated, so as to determine the candidate layout unit area based on the feature similarity with the preset layout features, and then further determine the target layout unit based on the semantic relevance between the candidate layout unit and the preset keyword, and finally determine the signature area according to the position overlap degree between the target text area and the target layout unit. Therefore, the signature area positioning scheme of the embodiments of the present application greatly improves the accuracy of signature area positioning by combining the target area determined by text parsing and the target area determined by image processing, enables automatic positioning of the signature area, and improves the positioning efficiency of the signature area.
[0062] Embodiment 2
[0063] Figure 2 This is a schematic structural diagram of an embodiment of the signature area positioning device in the contract document provided by the present application. As Figure 2 shown in the figure, the signature area positioning device in the contract document provided by the embodiments of the present application may include: an acquisition module 21, a text parsing module 22, an image processing module 23, and a determination module 24.
[0064] The acquisition module 21 can be used to acquire a target contract document. In the embodiments of the present application, the acquisition module 21 can acquire a specified target contract document through a server. The target contract documents in the embodiments of the present application can cover various types, such as commercial contracts, labor contracts, lease contracts, sales contracts, etc. It should be noted that with the increasing complexity of contract services, a contract may involve multiple parties, such as Party A, Party B, Party C, etc. in a commercial contract. In view of this situation, in the signature area positioning method of the embodiments of the present application, the historical contract processing data can be read and the relevant contract documents can be parsed, or the contract documents pre-stored in a separate server or a server on which the method of the embodiments of the present application is deployed can be acquired. In this way, the acquisition module 21 can efficiently parse or directly acquire the target contract document, providing a basis for subsequent signature area positioning.
[0065] The text parsing module 22 can be used to calculate the first semantic relevance between each text unit in the target contract document and a preset keyword using a first machine learning model. The text units with the calculated first semantic relevance greater than the preset semantic threshold are determined as target text units; the target text areas corresponding to the target text units are determined according to the text unit position information of the target text units. For example, the text parsing module 22 can use the first machine learning model to parse the target contract document acquired by the acquisition module 21. For example, in the embodiments of the application, a semantic-based machine learning model can be used to parse the target contract document to obtain the word vectors of each word, and then, for example, a word vector model can be used to calculate the semantic similarity between each word vector and the preset keyword. In particular, in the embodiments of the present application, the preset keyword can be a word related to signing and / or stamping.
[0066] Then, the text parsing module 22 can sort the calculated first semantic relevance between each text unit and the preset keyword, and determine the text units with the semantic relevance greater than the preset semantic threshold as target text units. For example, in the embodiments of the present application, the text parsing module 22 can sort all text units in descending order according to the calculated semantic relevance, so that the text units greater than the preset semantic threshold are regarded as the text units most closely semantically associated with the preset keyword, and thus determined as target text units. In the embodiments of the present application, these target text units may contain key information related to contract signature, such as words containing "sign", "seal", "signing party", etc. or words semantically similar.
[0067] The text parsing module 22 can obtain the position information according to the determined target text unit. For example, it can obtain the coordinate information of the first text unit or the last text unit of the target text unit as its position information, so that the area range where a signature seal may exist can be further determined according to the target text unit. For example, in the embodiment of the present application, the target text unit may be a text unit containing the words "signature" or "seal", and therefore, the area where the signature seal finally exists is likely to be near this text unit. For example, usually, the user will sign or seal in the area on the right side or below the area marked with the words "signature" or "seal" in the contract. Therefore, when the target text unit containing the words "signature" or "seal" is determined, the text parsing module 22 can extend based on the coordinate of the first word or the last word in the determined target text unit to form the target text area.
[0068] For example. It can start from the coordinates of the first word or the last word of the target text unit, and can extend the first distance and the second distance respectively in the first direction and the second direction to form two adjacent sides of a rectangle, and correspondingly continue to use the first distance and the second distance to form the other two sides of the rectangle, so as to form a rectangular area with the target text unit as the vertex as the target text area. For example, usually, the writing direction of text is from left to right, and the arrangement direction of text is from top to bottom. Therefore, the first direction can be the right side direction on the page of the target contract document, and the second direction can be the lower side direction on the page of the target contract document, so that the lower right side of the target text unit can be used as the target text area. Of course, according to different written texts, the text direction of the contract document can also be different, and in the embodiment of the present application, the above text direction can be determined by parsing the target text unit of the contract document. For example, according to the semantic vector of the target text unit calculated by the first machine learning model, the semantic arrangement order of the words contained in the target text unit can be determined. For example, when the target text unit contains the two words "signature", the semantic arrangement order of the two words "sign" and "ature" contained therein can be determined according to the semantic vector, that is, the line direction of the words is from left to right, as the first direction.
[0069] Then, the second machine learning model can be used to calculate the semantic logical relationship between the target text unit and other line text units or other paragraph text units to determine the second direction. That is, based on the semantic logical relationship between the target text unit and the text units in the adjacent line or the text units in the adjacent paragraph, the arrangement relationship between different lines / paragraphs in the target contract document can be determined, so that the second direction can be determined. Then, the text parsing module 22 can also determine the value of the first distance according to the position information of adjacent characters in the target text unit. For example, the character distance can be calculated according to the coordinates of adjacent characters, that is, the length of each character, and then the value obtained by multiplying the length of each character by a first multiple, such as 10 characters, is used as the value of the first distance; the line distance between the target text unit and other line text units and the paragraph distance between the target text unit and other paragraph text units are calculated according to the position information of the target text unit, line text unit, and paragraph text unit, and the maximum value of the line distance and the paragraph distance is used as the value of the second distance. Therefore, a rectangular area can be determined that starts from the target text unit as the starting vertex and extends the first distance in the text arrangement direction, that is, the line direction, and extends the second distance in the paragraph arrangement direction. In other words, in the embodiments of the present application, the text unit position can be used as the starting vertex of the rectangular area, and the first distance and the second distance are extended from the text unit position in the first direction and the second direction perpendicular to the first direction respectively as the two adjacent sides of the rectangular area, thereby defining the rectangular area.
[0070] In addition, in the embodiments of the present application, the shape of the target text area is not limited to a rectangular shape, and other shapes can also be used. For example, the target text unit can be used as the origin, and a circular area can be defined with the first distance or the second distance as the radius as the target text area.
[0071] The image processing module 23 can be used to extract multiple image features of the target contract document using the second machine learning model; generate the position features of each image feature in the multiple image features relative to other image features based on the position relationship between the multiple image features; calculate the feature similarity between each position feature and each preset layout feature; determine the image features corresponding to the position features whose calculated feature similarity is greater than the preset layout threshold corresponding to each preset layout feature as the candidate layout unit areas; calculate the second semantic relevance between the classification information of the preset layout unit areas corresponding to each candidate layout unit and the preset keywords; and determine the candidate layout unit areas with the calculated second semantic relevance greater than the preset semantic threshold as the target layout unit areas.
[0072] For example, the image processing module 23 can use, for example, a convolutional neural network model to extract multiple image features from the target contract document. Then, in step S106, a convolutional neural network can also be used to determine the positional relationship between the extracted multiple image features, thereby generating positional features of each image feature relative to other image features.
[0073] Before extracting the image features, the target contract document can also be preprocessed first. For example, grayscale processing can be performed first to convert the color image into a grayscale image, which can reduce the data processing volume in subsequent steps. In addition, a filtering algorithm, such as Gaussian filtering, can be adopted to remove the noise in the converted grayscale image, making the image clearer and facilitating subsequent processing. In addition, the contrast of the converted grayscale image can be further enhanced through image enhancement techniques, such as histogram equalization, to highlight the features of signatures and seals.
[0074] The image processing module 23 can calculate the feature similarity between the generated positional features and the preset layout features. Then, based on the calculated feature similarity, candidate layout units can be further determined. For example, the image processing module 23 can first sort the similarities, and further obtain the position information of the image features greater than the preset threshold, and determine the region containing a predetermined number or more of such image features greater than the preset threshold as the candidate layout unit region.
[0075] In the embodiment of the present application, the candidate layout unit region is a preset layout unit region corresponding to a certain preset layout feature. For example, in the embodiment of the present application, preset layout unit regions such as a title region, a body region, and a signature region can be preset. The image processing module 23 can calculate the similarity between the positional features of each image feature and the preset preset layout unit regions to determine which layout unit the image feature belongs to. Then, the image features with similarities greater than the preset threshold can be aggregated, and the coordinates of the outermost image features among these image features can be selected as the boundaries of the candidate layout unit region based on the position information of these image features to generate the candidate layout unit region. In addition, in the embodiment of the present application, the region center of these image features can also be used as the region center according to the region information of the preset layout unit regions corresponding to these image features, such as the region size information, and the region size information of the corresponding preset layout unit region can be used as the size of the candidate layout unit region to generate the candidate layout unit region.
[0076] The image processing module 23 can obtain the classification information for the preset layout unit to which each image feature belongs. For example, it belongs to one of the classifications such as the title area, the body text area, the signature area, etc. Then, it can calculate the semantic relevance between the obtained classification information and the preset keywords used by the text parsing module 22, that is, whether the classification corresponding to the image feature determined by the image processing module 23 is closely related to the preset keywords. For example, if the image feature determined by the image processing module 23 belongs to the title area, and the preset keyword used by the text parsing module 22 is "signature", then the image processing module 23 can calculate the semantic relevance between the classification "title" to which the image feature belongs and "signature".
[0077] The determination module 24 can be used to calculate the position overlap degree between the target text area and the target layout unit area; and determine the target text area with the calculated position overlap degree greater than the preset overlap threshold as the signature area.
[0078] The determination module 24 can sort based on the semantic relevance between the classification information to which each calculated image feature belongs and the preset keywords, and determine the candidate layout unit area with a second semantic relevance greater than the preset semantic threshold as the target layout unit area. Then, it can calculate the position overlap degree between the target layout unit area determined by the image processing module 23 and the target text area determined by the text parsing module 22. For example, the position overlap degree can be calculated based on the boundary coordinate information of the target layout unit area and the boundary coordinate information of the target text area, and the determination module 24 can determine the target text area with an overlap degree greater than the preset threshold as the signature area. That is, the target text area determined based on text parsing and the target layout unit area determined based on image processing are combined. When the overlap degree of the areas determined by the two methods is relatively high, it can be confirmed that the possibility of this area belonging to the signature area is also relatively high. Therefore, it can avoid the misrecognition caused by determining the signature area by the text parsing method or the image processing method, thereby improving the accuracy of signature area recognition.
[0079] In addition, in the embodiment of the present application, when determining the target layout unit area based on the candidate layout unit area with the calculated second semantic relevance greater than the preset semantic threshold, the image features contained therein can be extracted from the candidate layout unit area with the calculated second semantic relevance greater than the preset semantic threshold using a convolutional neural network, and then the feature matching degree between the extracted image features and the preset keyword features can be calculated; and the candidate layout unit containing the image features with a feature matching degree greater than the preset matching threshold is determined as the target layout unit.
[0080] In addition, in the embodiment of the present application, after the determination module 24 determines the signature area, it may further use a convolutional neural network to extract the image features contained therein for the determined signature area; calculate the feature matching degree between the extracted image features and the preset keyword features; determine the image features with the feature matching degree greater than the preset matching threshold as candidate key features; calculate the position information of the edge pixels of the candidate key features; and update the boundary of the signature area determined in step S112 according to the position information.
[0081] According to the signature area positioning device in the contract document of the embodiment of the present application, by performing text parsing on the target contract document to determine the target text area related to the preset keyword, then using the second machine learning model to extract the image features of the target contract document, and calculating the position features of the image features relative to other image features, so as to determine the candidate layout unit area based on the feature similarity with the preset layout features, and then further determine the target layout unit based on the semantic relevance between the candidate layout unit and the preset keyword, and finally determine the signature area according to the position overlap degree between the target text area and the target layout unit. Therefore, the signature area positioning solution of the embodiment of the present application greatly improves the accuracy of signature area positioning by combining the target area determined by text parsing and the target area determined by image processing, enabling automatic positioning of the signature area and improving the positioning efficiency of the signature area.
[0082] Embodiment III
[0083] The internal functions and structures of the signature area positioning device in the contract document are described above, and the device can be implemented as an electronic device. Figure 3 It is a schematic structural diagram of an embodiment of the electronic device provided by the present application. As Figure 3 shown, the electronic device includes a memory 31 and a processor 32.
[0084] The memory 31 is used to store programs. In addition to the above programs, the memory 31 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0085] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0086] The processor 32 is not limited to a central processing unit (CPU), but may also be a processing chip such as a graphics processing unit (GPU), a field programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. The processor 32 is coupled to the memory 31 and executes the program stored in the memory 31. When the program runs, it executes the signature area positioning method in the contract document of the first embodiment above.
[0087] Further, as Figure 3 shown, the electronic device may further include other components such as a communication component 33, a power supply component 34, an audio component 35, and a display 36. Figure 3 Only some components are schematically shown in Figure 3 the figure, which does not mean that the electronic device only includes
[0088] The communication component 33 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on communication standards, such as WiFi, 3G, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 33 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 33 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0089] The power supply component 34 provides power for various components of the electronic device. The power supply component 34 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device.
[0090] The audio component 35 is configured to output and / or input audio signals. For example, the audio component 35 includes a microphone (MIC). When the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory 31 or transmitted via the communication component 33. In some embodiments, the audio component 35 further includes a speaker for outputting audio signals.
[0091] The display 36 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can detect not only the boundaries of touch or swipe actions, but also the duration and pressure associated with the touch or swipe operations.
[0092] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0093] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for locating a signature area in a contract document, characterized in that: include: Obtain target contract documents; Using a first machine learning model, calculating a first semantic relevance between each text unit in the target contract document and a preset keyword, wherein the preset keyword is a word related to signature and / or seal; Determine a text unit whose calculated first semantic relevance is greater than a preset semantic threshold as a target text unit; Determining a target text area corresponding to each target text unit according to the text unit position information of the target text unit, wherein the target text area is a rectangular area, wherein the rectangular area has the text unit position as a starting vertex, and a first distance and a second distance extending from the text unit position toward a first direction and a second direction perpendicular to the first direction respectively as two adjacent sides of the rectangular area, thereby defining the rectangular area; Using a second machine learning model, extracting a plurality of image features of the target contract document; generating a position feature of each image feature in the plurality of image features relative to other image features based on the positional relationship between the plurality of image features; Calculating feature similarity between each position feature and each preset layout feature; Determine a candidate layout unit region based on the image feature corresponding to the position feature whose calculated feature similarity is greater than the preset layout threshold corresponding to each preset layout feature, wherein the candidate layout unit region is the preset layout unit region corresponding to the preset layout feature; Calculating a second semantic relevance between classification information of a preset layout unit area corresponding to each candidate layout unit area and the preset keyword; Determine the candidate layout unit region whose calculated second semantic relevance is greater than a preset semantic threshold as a target layout unit region; Calculating the position overlap between the target text area and the target layout unit area; The target text area whose calculated position overlap is greater than a preset overlap threshold is determined as the signature area.
2. The method for locating the signature area in a contract document according to claim 1, characterized in that: The step of determining the target text area corresponding to each target text unit according to the text unit position information of the target text unit comprises: Determining a semantic arrangement order of characters contained in the target text unit according to the semantic vector of the target text unit calculated by the first machine learning model; Using a second machine learning model to calculate the semantic logical relationship between the target text unit and other line text units or other paragraph text units, wherein the other line text units are text units in lines adjacent to the line where the target text unit is located, and the other paragraph text units are text units in paragraphs adjacent to the paragraph where the target text unit is located; Determining the first direction and the second direction according to the semantic arrangement order and the semantic logical relationship, wherein the first direction is the line direction of the target text unit, and the second direction is the paragraph direction of the target text unit; Determine the value of the first distance according to the position information of the adjacent characters in the target text unit; The line distance between the target text unit and the other line text units and the segment distance between the target text unit and the other segment text units are calculated based on the position information of the target text unit, the other line text units and the other segment text units, and the maximum value of the line distance and the segment distance is used as the value of the second distance.
3. The method for locating the signature area in a contract document according to claim 2, characterized in that: The step of determining the value of the first distance according to the position information of adjacent characters in the target text unit comprises: Calculating the spacing between adjacent characters based on the adjacent character position information of the target text unit; A value obtained by multiplying the interval by a first multiple is taken as the value of the first distance.
4. The method for locating the signature area in a contract document according to claim 3, characterized in that: The step of determining the candidate layout unit region whose calculated second semantic relevance is greater than a preset semantic threshold as the target layout unit region comprises: For the candidate layout unit regions whose calculated second semantic relevance is greater than a preset semantic threshold, a convolutional neural network is used to extract the image features contained therein; Calculating the feature matching degree between the extracted image features and the preset keyword features; A candidate layout unit region including an image feature having a feature matching degree greater than a preset matching threshold is determined as a target layout unit.
5. The method for locating the signature area in a contract document according to claim 4, characterized in that: The method further comprises: Using a convolutional neural network to extract image features contained in the signature area; Calculating the feature matching degree between the extracted image features and the preset keyword features; An image feature whose feature matching degree is greater than a preset matching threshold is determined as a candidate key feature; Calculate the position information of the edge pixels of the candidate key features; The boundary of the signature area is updated according to the position information.
6. The method for locating the signature area in a contract document according to claim 1, characterized in that: The obtaining of the target contract document comprises: The target contract document in the form of a color image is grayscaled to obtain a grayscale image as the target contract document.
7. The method for locating the signature area in a contract document according to claim 1, characterized in that: The method further comprises: For the signature area, a coordinate calculation algorithm is used to calculate the position of the signature area in the target contract document.
8. The method for locating the signature area in a contract document according to claim 1, characterized in that: The method further comprises: Performing semantic analysis on the target text unit to obtain document location information for the preset keyword; Determine a candidate target area in the target contract document according to the document location information; And the extracting of multiple image features of the target contract document using the second machine learning model includes: Using a second machine learning model, multiple image features of the candidate target area are extracted.
9. A device for locating a signature area in a contract document, characterized in that: include: An acquisition module, used to acquire a target contract document; A text parsing module, configured to use a first machine learning model to calculate a first semantic relevance between each text unit in the target contract document and a preset keyword, wherein the preset keyword is a word related to signature and / or stamping; determine a text unit whose calculated first semantic relevance is greater than a preset semantic threshold as a target text unit; determine a target text area corresponding to each target text unit according to text unit position information of the target text unit, wherein the target text area is a rectangular area, wherein the rectangular area has the text unit position as a starting vertex, and has a first distance and a second distance extending from the text unit position toward a first direction and a second direction perpendicular to the first direction as two adjacent sides of the rectangular area, thereby defining the rectangular area; An image processing module is used to extract multiple image features of the target contract document using a second machine learning model; generate position features of each image feature in the multiple image features relative to other image features based on the positional relationship between the multiple image features; calculate the feature similarity between each position feature and each preset layout feature; determine the image feature corresponding to the position feature whose calculated feature similarity is greater than the preset layout threshold corresponding to each preset layout feature as a candidate layout unit area, wherein the candidate layout unit area is the preset layout unit area corresponding to the preset layout feature; calculate the second semantic relevance between the classification information of the preset layout unit area corresponding to each candidate layout unit area and the preset keyword; determine the candidate layout unit area whose calculated second semantic relevance is greater than the preset semantic threshold as the target layout unit area; The determination module is used to calculate the position overlap between the target text area and the target layout unit area; and determine the target text area whose calculated position overlap is greater than a preset overlap threshold as the signature area.
10. An electronic device, characterized in that: include: Memory, used to store programs; A processor, configured to run the program stored in the memory to execute the method for locating a signature area in a contract document as described in any one of claims 1 to 8.