Pictographic phrase paraphrase mining method and device, equipment, medium and product
Through the combination of the Transformer model and the graphic alignment model, the probability distribution between the characters of pictograms is determined and the word participle in the descriptive text is matched, which solves the problem of digging the meaning of pictogram phrases, and realizes the effective and accurate digging of the definition of object-type phrases.
Patent Information
- Application Number
- CN202510536901.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to effectively explore the meaning of phrases in hieroglyphs, which leads to difficulties in researching and inheriting the meaning of hieroglyphs.
The inter-word probability distribution of pictograms in the corpus is determined through the Transformer model, the candidate phrases are determined, and the participle of the candidate phrases is matched from the corresponding single-sentence interpretation text through the graphic and text alignment model to obtain the meaning of the candidate phrases.
It realizes effective and accurate exploration of the definition of object morphology phrases, and provides important corpus-supported research on hieroglyphs and languages.
Smart Images

Figure CN120068854A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method, device, equipment, medium and product for mining the interpretation of pictographic character phrases. Background Technique
[0002] As an early stage of the origin of writing, pictographic characters are usually not languages that can be translated word by word, and their interpretations have great uncertainties. Usually, they can only be passed down in the form of oral instruction, but most of them have been lost so far. For example, Figure 1 As shown in the figure, Dongba script is a kind of pictographic character created by the ancestors of the Naxi ethnic group in China. The Dongba classics recorded in it are known as the "encyclopedia" of Naxi society and are listed as "Memory of the World". Currently, there are more than 30,000 volumes of Dongba classics in circulation, a total of more than 1,400 kinds. However, it is facing the endangered situation of no one being able to interpret it.
[0003] A phrase is a combination of characters that appears repeatedly and has a fixed meaning. Taking Dongba script as an example, although no scholar has proposed the existence of phrases in Dongba script so far, in the annotated Dongba ancient books, it can be found that there are indeed combinations of characters that appear repeatedly and have fixed Chinese interpretations: for example, the combination of "day" and "wooden barrel" means "east"; the combination of "day" and "egg" means "west", etc.
[0004] It can be seen that how to effectively mine the interpretations of phrases in pictographic characters not only has extremely high academic value for the research of pictographic characters and languages, but also can provide important language materials for the research of pictographic characters and their interpretations. Summary of the Invention
[0005] The present application provides a method, device, equipment, medium and product for mining the interpretation of pictographic character phrases to effectively and accurately mine the interpretations of phrases in pictographic characters.
[0006] In a first aspect, an embodiment of the present application provides a method for mining the interpretation of pictographic character phrases, including:
[0007] Determine the inter-character probability distribution of pictographic characters in the corpus through a Transformer model, and determine candidate phrases according to the inter-character probability distribution;
[0008] For each of the candidate phrases, determine the corresponding image segment of the candidate phrase and the single-sentence interpretation text corresponding to the candidate phrase;
[0009] Through a graphic-text alignment model, determine the word segmentation corresponding to the candidate phrase in the single-sentence interpretation text to obtain the interpretation corresponding to the candidate phrase.
[0010] In a second aspect, an embodiment of the present application further provides a device for mining the interpretation of pictographic character phrases, including:
[0011] A phrase determination module, configured to determine the inter-character probability distribution of hieroglyphs in a corpus through a Transformer model, and determine candidate phrases according to the inter-character probability distribution;
[0012] A phrase processing module, configured to, for each of the candidate phrases, determine the image segment corresponding to the candidate phrase and the single-sentence paraphrase text corresponding to the candidate phrase;
[0013] An interpretation determination module, configured to determine the word segmentation corresponding to the candidate phrase in the single-sentence paraphrase text through a text-image alignment model, and obtain the interpretation corresponding to the candidate phrase.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0015] One or more processors;
[0016] A storage device, configured to store one or more programs;
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for mining the interpretation of hieroglyphic phrases as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for mining the interpretation of hieroglyphic phrases as described in the first aspect is implemented.
[0019] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program and / or instructions, and when the computer program and / or instructions are executed by a processor, the method for mining the interpretation of hieroglyphic phrases as described in any of the above embodiments is implemented.
[0020] An embodiment of the present application provides a method, device, equipment, medium and product for mining the interpretation of hieroglyphic phrases. The method for mining the interpretation of hieroglyphic phrases includes: determining the inter-character probability distribution of hieroglyphs in a corpus through a Transformer model, and determining candidate phrases according to the inter-character probability distribution; for each of the candidate phrases, determining the image segment corresponding to the candidate phrase and the single-sentence paraphrase text corresponding to the candidate phrase; determining the word segmentation corresponding to the candidate phrase in the single-sentence paraphrase text through a text-image alignment model, and obtaining the interpretation corresponding to the candidate phrase. The above technical solution determines candidate phrases according to the inter-character probability distribution, and uses a text-image alignment model to match the word segmentation corresponding to the candidate phrase from the corresponding single-sentence paraphrase text, and obtains the interpretation of the candidate phrase, realizing effective and accurate mining of the interpretation of hieroglyphic phrases. Description of the Drawings
[0021] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.
[0022] Figure 1 A schematic diagram of Dongba characters provided for this application;
[0023] Figure 2 A flowchart of a method for mining the interpretation of pictographic character phrases provided in an embodiment of this application;
[0024] Figure 3 A schematic diagram of a corpus provided for an embodiment;
[0025] Figure 4 A schematic diagram of a Transformer model provided for an embodiment;
[0026] Figure 5 A schematic diagram of an inference process for determining candidate phrases using a Transformer model provided for an embodiment;
[0027] Figure 6 A schematic diagram of a process for aligning text and images provided for an embodiment;
[0028] Figure 7 A schematic diagram of a process for mining phrases of pictographic characters provided for an embodiment;
[0029] Figure 8 A schematic structural diagram of a device for mining the interpretation of pictographic character phrases provided in an embodiment of this application;
[0030] Figure 9 A schematic structural diagram of an electronic device provided in an embodiment of this application. Specific Embodiments
[0031] The following further describes this application in detail with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining this application and not for limiting this application. Additionally, it should be noted that for ease of description, only parts related to this application rather than all structures are shown in the drawings.
[0032] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0033] It should be noted that the concepts such as "first" and "second" mentioned in the embodiments of the present application are only used to distinguish different devices, modules, units, or other objects, and are not used to limit the order or interdependence of the functions performed by these devices, modules, units, or other objects.
[0034] In addition, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0035] In the technical solution of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0036] It should be noted that in the embodiments of the present application, some existing industry solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used the relevant content of this solution.
[0037] Figure 2 It is a flowchart of a method for mining the interpretation of hieroglyphic word groups provided for the embodiments of the present application. This embodiment is applicable to the situation of determining the interpretation of hieroglyphic word groups. Specifically, the method for mining the interpretation of hieroglyphic word groups can be executed by a device for mining the interpretation of hieroglyphic word groups. The device for mining the interpretation of hieroglyphic word groups can be implemented in software and / or hardware and integrated in an electronic device. The electronic device includes, but is not limited to, devices with computing functions such as a computer, a smart phone, or a server.
[0038] As Figure 2 shown, the method specifically includes the following steps:
[0039] S110. Determine the inter-character probability distribution of hieroglyphs in the corpus through a Transformer model, and determine candidate word groups according to the inter-character probability distribution.
[0040] In this embodiment, the corpus can be understood as a collection of hieroglyphics, which contains a large number of images of hieroglyphics. These hieroglyphics can be sourced from paper or electronic ancient books of hieroglyphics, or obtained through channels such as the Internet, and can be recorded in the form of single sentences or paragraphs. In addition, these hieroglyphics have corresponding paraphrase texts, that is, the results after being translated into modern natural language (such as Chinese). Figure 3 FIG. is a schematic diagram of a corpus provided for an embodiment. As Figure 3 shown, each of the images below can represent a single sentence of Dongba script. There are m images of single-sentence Dongba script in total, and their corresponding m Chinese translation texts (that is, the paraphrase texts of the single-sentence Dongba script) are shown in the corresponding images above.
[0041] The inter-character probability distribution can be used to describe the correlation between hieroglyphics in the corpus, and can also be understood as the probability of the combination of hieroglyphics. The probability of combination is related to the number of times or frequency of the combination of hieroglyphics in the corpus. The greater the correlation or the probability of combination, the more likely the corresponding hieroglyphics are to form a phrase. A phrase can also be understood as a combination of adjacent characters. It can be understood that a phrase can include two or more hieroglyphics. For example, if the probability that a hieroglyphic (denoted as hieroglyphic A) is followed by another hieroglyphic (denoted as hieroglyphic B) is relatively high, then the combination of hieroglyphic A and hieroglyphic B can be used as a phrase; another example is that if the probability that the combination of hieroglyphic A and hieroglyphic B is followed by another hieroglyphic (denoted as hieroglyphic C) is relatively high, then the combination of hieroglyphic A, hieroglyphic B, and hieroglyphic C can be used as a phrase. Based on this, multiple determined phrases can be used as candidate phrases, that is, phrases to be further mined for paraphrases.
[0042] The Transformer model can be used to determine the inter-character probability distribution of hieroglyphics in the corpus. By learning the probability relationship between single characters, adjacent character combinations with greater correlation can be inferred. Exemplarily, the hieroglyphics in the corpus can be recognized. Assuming that there are x (x is a positive integer) images corresponding to hieroglyphics in the corpus, then x encodings can be obtained, and then the Transformer model is used to mine candidate phrases according to the inter-character probability distribution.
[0043] As an example, the Transformer model includes a decoder. The Transformer model can be composed of multiple Transformer modules, and the core Transformer module is composed of a masked multi-head self-attention mechanism and a feed-forward neural network.
[0044] S120. For each of the candidate phrases, determine the image segment corresponding to the candidate phrase and the single-sentence paraphrase text corresponding to the candidate phrase;
[0045] In this embodiment, an image segment can be understood as an image corresponding to a candidate phrase, and the image segment does not contain hieroglyphs other than the candidate phrase. For each candidate phrase, its corresponding image segment can be determined. For example, the corresponding image segment is intercepted from the image of the single sentence or paragraph where the candidate phrase is located. Correspondingly, the paraphrase text of the single sentence where the candidate phrase is located can be obtained, and the purpose is to extract the paraphrase corresponding to the candidate phrase from the single sentence paraphrase text through subsequent processing.
[0046] S130. Through a graphic-text alignment model, determine the word segmentation corresponding to the candidate phrase in the single sentence paraphrase text to obtain the paraphrase corresponding to the candidate phrase.
[0047] In this embodiment, the graphic-text alignment model can be understood as a model used to achieve the matching of images and texts according to visual features and language features. The graphic-text alignment model can be a vision-language model, such as a Contrastive Language-Image Pre-training (CLIP) model or a Bootstrapping Language-Image Pre-training (BLIP) model, etc. The single sentence paraphrase text usually includes multiple word segmentations. The graphic-text alignment model can be used to find the word segmentation corresponding to the candidate phrase from multiple word segmentations to achieve graphic-text alignment, so as to obtain the paraphrase corresponding to the candidate phrase. The process of graphic-text alignment can be understood as a process of cross-modal alignment between hieroglyphs and paraphrases. In this process, the consistency of the candidate phrase paraphrase can be verified.
[0048] A method for mining the paraphrase of a hieroglyph phrase provided by an embodiment of this application can convert the images of each hieroglyph in a corpus into corresponding codes through hieroglyph recognition. The Transformer including a decoder can determine candidate phrases according to the inter-character probability distribution, and then use a graphic-text alignment model to match the word segmentation corresponding to the candidate phrase from the corresponding single sentence paraphrase text to obtain the paraphrase of the candidate phrase, realizing effective and accurate mining of the paraphrase of the hieroglyph phrase.
[0049] In one embodiment, the Transformer model includes a masked multi-head self-attention mechanism and a feed-forward neural network; the Transformer model is used to learn the context relevance features of hieroglyphs based on the causal masking strategy in the masked self-attention mechanism and output the inter-character probability between different hieroglyphs.
[0050] Figure 4 Schematic diagram of a Transformer model provided for one embodiment. As Figure 4As shown in the figure, the Transformer model may include multiple layers of Transformer modules. The core Transformer module consists of a masked multi-head self-attention mechanism and a feed-forward neural network. By training on the Transformer model, the context information and inter-character correlations of hieroglyphs are modeled. Under the condition of unlabeled data, this model learns the context correlation features of hieroglyphs and outputs the probability distribution between single characters. By inferring the probabilities between characters, adjacent characters with closely related semantics are combined to generate candidate phrases.
[0051] Taking the hieroglyph encoding sequence as the input, the autoregressive generation process is achieved through the causal masking strategy in the masked self-attention mechanism. Exemplarily, given a starting sequence , the Transformer model can calculate the probability distribution of the next based on the previous sequence:
[0052] ;
[0053] where, represents the context embedding generated by the Transformer model. Based on the causal masking mechanism, this embedding only contains the information of the sequence { }. and are the linear transformation parameters of the model, used to map the context embedding vector output by the Transformer to the vocabulary space. The softmax function converts this mapping into a probability distribution for predicting the next hieroglyph . On this basis, the phrases of hieroglyphs can be comprehensively and accurately recognized.
[0054] In one embodiment, each candidate phrase includes x hieroglyphs, where x is an integer greater than or equal to 2. The determination of candidate phrases according to the inter-character probability distribution includes: if the inter-character probability distribution of any x ordered hieroglyphs is greater than a set threshold, then the x ordered hieroglyphs are used as a candidate phrase.
[0055] In this embodiment, the candidate phrases can be composed of two or more closely related hieroglyphs combined. Figure 5 is a schematic diagram of the inference process for determining candidate phrases using the Transformer model provided in an embodiment. As shown in Figure 5 As shown, given a hieroglyph (encoded as y1), the probability distribution between this hieroglyph and the rest of the hieroglyphs in the corpus can be calculated. If the probability of any hieroglyph appearing after this hieroglyph is relatively large (greater than the set threshold), then the relationship between this hieroglyph and this other hieroglyph is relatively close, and the combination of this hieroglyph and this other hieroglyph can be used as a candidate phrase. Based on this, all two-character candidate phrases in the corpus can be determined. Further, given a combination of two adjacent hieroglyphs (encoded as y1, y2), the probability distribution between these two adjacent and orderly arranged hieroglyphs and the rest of the hieroglyphs in the corpus can be calculated. If the probability of any hieroglyph appearing after these two adjacent hieroglyphs is relatively large, then the relationship between these two adjacent hieroglyphs and this hieroglyph is relatively close, and the combination of these two adjacent hieroglyphs and this hieroglyph can be used as a candidate phrase. Based on this, all three-character candidate phrases in the corpus can be determined, and so on, all candidate phrases of any number of characters in the corpus can be determined.
[0056] In one embodiment, in addition to determining whether a combination can be used as a candidate phrase based on whether the inter-character probability distribution is greater than the set threshold, another method can also be used: calculate the probability distribution between a hieroglyph and the rest of the hieroglyphs in the corpus, and select a set number of hieroglyphs that are most closely related to this hieroglyph according to the probability distribution. Combine this hieroglyph with each of the selected hieroglyphs as candidate phrases. For example, there are 1500 hieroglyphs in the corpus. After a hieroglyph is given, calculate the inter-character probability distribution between this hieroglyph and the remaining 1499 hieroglyphs, and select the 100 hieroglyphs with the highest probability. Combine the given hieroglyph with each of these 100 hieroglyphs to obtain 100 two-character candidate phrases. Similarly, calculate the probability distribution between two adjacent hieroglyphs and the rest of the hieroglyphs in the corpus, select a set number of hieroglyphs that are most closely related to these two adjacent hieroglyphs, and combine these two adjacent hieroglyphs with each of the selected hieroglyphs as candidate phrases. For example, there are 1500 hieroglyphs in the corpus. After two hieroglyphs are given, calculate the inter-character probability distribution between these two hieroglyphs and the remaining 1498 hieroglyphs, and select the 100 hieroglyphs with the highest probability. Combine the given two adjacent hieroglyphs with each of these 100 hieroglyphs to obtain 100 three-character candidate phrases. And so on, all candidate phrases of any number of characters in the corpus can be determined.
[0057] In one embodiment, for each candidate phrase, the corresponding single-sentence paraphrase text includes at least two word segments; the method of determining the word segments in the single-sentence paraphrase text corresponding to the image segment through the graphic-text alignment model to obtain the paraphrase corresponding to the candidate phrase includes:
[0058] S1310. Calculate the feature similarity of the image segment with respect to each word segment through the graphic-text alignment model;
[0059] S1320. Take the word segment with the highest feature similarity as the corresponding interpretation of the candidate phrase;
[0060] Among them, the graphic-text alignment model is trained based on multiple groups of sample data, and each group of sample data includes a single-sentence image of hieroglyphics and the corresponding single-sentence interpretation text.
[0061] Exemplarily, for any candidate phrase, perform feature encoding on its image segment, and perform feature encoding on all word segments in its corresponding single-sentence interpretation text. Then, using the graphic-text alignment model, the feature similarity of the image segment with respect to each word segment can be calculated according to the feature encoding, and further, the word segment with the highest feature similarity can be taken as the corresponding interpretation of the candidate phrase. On this basis, the consistency of the graphic-text matching of the candidate phrase can be ensured, and accurate phrase mining can be realized.
[0062] In one embodiment, the step of taking the word segment with the highest feature similarity as the corresponding interpretation of the candidate phrase includes: if the image segment appears at least twice, and the word segment with the highest feature similarity corresponding to each appearance of the image segment is the same, then take the word segment as the corresponding interpretation of the candidate phrase.
[0063] Exemplarily, by using the graphic-text alignment model to find the word segment with the highest feature similarity corresponding to each candidate phrase in the corpus, it can be understood that the graphic-text alignment result of each candidate phrase is obtained. Integrate and screen the alignment results of each candidate phrase. If a candidate phrase appears twice or more and the word segments in each alignment result are the same, the consistency of the interpretation of the candidate phrase can be verified, and thus it can be determined that the candidate phrase is a phrase of hieroglyphics with a clear interpretation. On this basis, the reliability and accuracy of phrase determination can be improved.
[0064] In one embodiment, the step of determining the image segment corresponding to the candidate phrase and the single-sentence interpretation text corresponding to the candidate phrase includes:
[0065] S1210. Retrieve the single-sentence image corresponding to the candidate phrase and the single-sentence interpretation text corresponding to the single-sentence image from the corpus;
[0066] S1230. Segment the single-sentence image to obtain the image segment of the candidate phrase.
[0067] Before determining the word segment corresponding to the candidate phrase in the single-sentence interpretation text through the graphic-text alignment model, it further includes:
[0068] S1230. Tokenize the single-sentence paraphrased text.
[0069] Figure 6 A schematic diagram of a graphic-text alignment process provided for an embodiment. As Figure 6 shown, taking the CLIP model as an example of the graphic-text alignment model, the CLIP model can be pre-trained using paired single-sentence images of hieroglyphics (taking Dongba script as an example) and their corresponding single-sentence paraphrased texts. The CLIP model can simultaneously process the feature vectors of images and texts, learn and establish the semantic relationship between images and texts; retrieve the positions of each candidate phrase in the corpus, and extract the image segments of the candidate phrases from the single-sentence images; in addition, tokenize the Chinese paraphrased text of the whole sentence or whole paragraph of hieroglyphics, and the obtained (x) tokens are subsequently used to match the image segments of the candidate phrases; in the graphic-text alignment (semantic matching) stage, using the pre-trained CLIP model, by calculating the feature similarity between the candidate phrase image embedding (feature encoding of the image segment) and the Chinese text embedding (feature encoding of the token), automatically match the candidate phrase to the Chinese paraphrased token with the highest semantic consistency (for example, for token 3), and after screening and determination, finally generate a hieroglyphic phrase with clear semantics.
[0070] Figure 7 A schematic diagram of a process for mining hieroglyphic phrases provided for an embodiment. As Figure 7 shown, this process makes full use of the data mining and semantic analysis capabilities of artificial intelligence technology. Taking Dongba characters as an example, in the first stage, using the Transformer model, according to the inter-character probability distribution, the combinations of adjacent hieroglyphics with close relationships (such as those that appear repeatedly) can be accurately and comprehensively found in the corpus; in the second stage, using the phrase graphic-text matching process of cross-modal semantic alignment, further verify the consistency of the Chinese paraphrases of these combinations of adjacent hieroglyphics, so as to extract the combinations of adjacent hieroglyphics that appear repeatedly and have fixed paraphrases in hieroglyphics, realizing the effective and accurate mining of the paraphrases of hieroglyphic phrases.
[0071] Figure 8 A schematic structural diagram of a device for mining the paraphrases of hieroglyphic phrases provided for an embodiment of the present application. The device for mining the paraphrases of hieroglyphic phrases provided in this embodiment includes:
[0072] A phrase determination module 210, configured to determine the inter-character probability distribution of hieroglyphics in the corpus through the Transformer model, and determine candidate phrases according to the inter-character probability distribution;
[0073] A phrase processing module 220, configured to, for each of the candidate phrases, determine the image segment corresponding to the candidate phrase and the single-sentence paraphrased text corresponding to the candidate phrase;
[0074] An interpretation determination module 230 is configured to determine the word segmentation corresponding to the candidate phrase in the single-sentence interpretation text through a graphic-text alignment model, so as to obtain the interpretation corresponding to the candidate phrase.
[0075] The device determines candidate phrases according to the inter-character probability distribution, and uses a graphic-text alignment model to match the word segmentation corresponding to the candidate phrases from the corresponding single-sentence interpretation text, so as to obtain the interpretations of the candidate phrases, realizing effective and accurate mining of the interpretations of hieroglyphic phrases.
[0076] Based on any of the above embodiments, each candidate phrase includes x hieroglyphs, where x is an integer greater than or equal to 2; the determining the candidate phrase according to the inter-character probability distribution includes: if the inter-character probability distribution of any x ordered hieroglyphs is greater than a set threshold, then taking the x ordered hieroglyphs as a candidate phrase.
[0077] Based on any of the above embodiments, for each candidate phrase, the single-sentence interpretation text includes at least two word segmentations;
[0078] The interpretation determination module 230 includes:
[0079] A calculation unit is configured to calculate the feature similarity between the image segment and each word segmentation through a graphic-text alignment model;
[0080] A determination unit is configured to take the word segmentation with the highest feature similarity as the interpretation corresponding to the candidate phrase;
[0081] Wherein, the graphic-text alignment model is trained based on multiple groups of sample data, and each group of sample data includes a single-sentence image of hieroglyphs and the corresponding single-sentence interpretation text.
[0082] Based on any of the above embodiments, the determination unit is specifically configured to: if the image segment appears at least twice, and the word segmentations with the highest feature similarity corresponding to each appearance of the image segment are the same, then taking the word segmentations as the interpretations corresponding to the candidate phrases.
[0083] Based on any of the above embodiments, the phrase processing module 220 includes:
[0084] A retrieval unit is configured to retrieve the single-sentence image corresponding to the candidate phrase and the single-sentence interpretation text corresponding to the single-sentence image from a corpus;
[0085] A segmentation unit is configured to segment the single-sentence image to obtain the image segment of the candidate phrase.
[0086] The device further includes a word segmentation unit, configured to perform word segmentation on the single-sentence paraphrase text before determining, through the graphic-text alignment model, the word segmentation corresponding to the candidate phrase in the single-sentence paraphrase text.
[0087] The device for mining the paraphrase of hieroglyphic phrases provided by the embodiment of the present application can be used to execute the method for mining the paraphrase of hieroglyphic phrases provided by any of the above embodiments, and has corresponding functions and beneficial effects.
[0088] Figure 9 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present application. The electronic device 10 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 10 can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, user equipment, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.
[0089] As Figure 9 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0090] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, wireless networks.
[0091] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above.
[0092] In some embodiments, the methods of the above embodiments may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the methods described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the methods of any of the above embodiments in any other suitable manner (e.g., by means of firmware).
[0093] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0094] The computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to the processors of general-purpose computers, special-purpose computers, or other programmable data processing devices, such that when the computer programs are executed by the processors, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0095] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0096] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device 10 having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device 10. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0097] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0098] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0099] The embodiments of the present application further provide a computer program product, including a computer program and / or instructions, which, when executed by a processor, implement the method for mining the interpretation of hieroglyphic phrases as described in any of the above embodiments.
[0100] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved, and no limitation is imposed herein.
[0101] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A method for mining the meaning of pictographic phrases, characterized in that: include: Determine the probability distribution of pictographic characters in the corpus through the Transformer model, and determine the candidate phrases based on the probability distribution of pictographic characters; For each candidate phrase, determining an image segment corresponding to the candidate phrase and a single sentence interpretation text corresponding to the candidate phrase; The word segments corresponding to the candidate phrase in the single sentence interpretation text are determined through the image-text alignment model to obtain the interpretation corresponding to the candidate phrase.
2. The method according to claim 1, characterized in that The Transformer model includes a masked multi-head self-attention mechanism and a feedforward neural network; The Transformer model is used to learn the contextual relevance features of pictograms based on the causal masking strategy in the masked self-attention mechanism, and output the inter-character probabilities between different pictograms.
3. The method according to claim 1, characterized in that Each candidate phrase includes x ideograms, where x is an integer greater than or equal to 2; The step of determining candidate phrases according to the probability distribution between characters includes: If the inter-character probability distribution of any x ordered pictograms is greater than a set threshold, the x ordered pictograms are taken as a candidate phrase.
4. The method according to claim 1, characterized in that For each of the candidate phrases, the single sentence interpretation text includes at least two participles; The step of determining the word segment corresponding to the image segment in the single sentence interpretation text by using the image-text alignment model to obtain the interpretation corresponding to the candidate phrase includes: Calculating the feature similarity of the image segment relative to each word segment through an image-text alignment model; The word with the highest feature similarity is used as the interpretation corresponding to the candidate phrase; The image-text alignment model is trained based on multiple sets of sample data, and each set of sample data includes a single-sentence image of a pictographic character and a corresponding single-sentence interpretation text.
5. The method according to claim 4, characterized in that The step of taking the segmented word with the highest feature similarity as the interpretation corresponding to the candidate phrase includes: If the image segment appears at least twice, and the segmented words with the highest feature similarity corresponding to the image segment are consistent each time, the segmented words are used as the interpretation corresponding to the candidate phrase.
6. The method according to claim 1, characterized in that The step of determining the image segment corresponding to the candidate phrase and the single sentence interpretation text corresponding to the candidate phrase includes: Retrieving from a corpus a single sentence image corresponding to the candidate phrase and a single sentence interpretation text corresponding to the single sentence image; Segmenting the single sentence image to obtain image segments of the candidate phrases; Before determining the word segment corresponding to the candidate phrase in the single sentence interpretation text through the image-text alignment model, the method further includes: The single sentence interpretation text is segmented.
7. A device for mining the meaning of pictographic phrases, characterized in that: include: A phrase determination module is used to determine the probability distribution between characters of pictographic characters in the corpus through the Transformer model, and determine candidate phrases according to the probability distribution between characters; A phrase processing module, for determining, for each candidate phrase, an image segment corresponding to the candidate phrase and a single sentence interpretation text corresponding to the candidate phrase; The interpretation determination module is used to determine the word segmentation corresponding to the candidate phrase in the single sentence interpretation text through a picture-text alignment model, and obtain the interpretation corresponding to the candidate phrase.
8. An electronic device, characterized in that: include: at least one processor; a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for mining the meaning of pictographic phrases according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for mining the meaning of pictographic phrases as described in any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program and / or instructions, characterized in that: When the computer program and / or instructions are executed by a processor, the method for mining the meaning of pictographic phrases as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Word paraphrase determination method and device, equipment and storage medium
CN111428721A
Dongba pictograph recognition method and device based on convolutional neural network
CN113837186A
Visual question and answer model and method based on multi-modal retrieval enhancement
CN119066174A
New word mining method, new word mining device and electronic equipment
CN119808779A