Text translation method, electronic device, storage medium and computer program product
By performing segmented processing and secondary translation of housing information pictures, combined with translation guidance and preset field dictionary adjustment, the problem of poor translation accuracy in the real estate field is solved, and more accurate and professional housing information translation is achieved, improving user experience.
Patent Information
- Application Number
- CN202510527457.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing text translation methods have poor translation accuracy in the real estate field, especially due to the lack of knowledge base support for the real estate field, which leads to mixed processing of professional fields and general content, resulting in mistranslation of terms or loss of formats, affecting user experience.
By identifying the original housing information text in the housing information picture, the text field is obtained by segmenting the process, and the first preset big model is used to translate it into the target language. Then, the second preset big model is used to detect whether the translation field meets the usage scenario of the target language, and the translation guidance is generated, and the translation results are corrected, and finally the translation results are adjusted according to the preset field dictionary.
It improves the accuracy and pertinence of translations, ensures that the translation results meet professional requirements in the real estate field, enhances the user experience, and provides more accurate and reliable translation results for housing information.
Smart Images

Figure CN120068895B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text translation, and in particular to a text translation method, system, electronic device, storage medium, and computer program product. Background Art
[0002] In the current real estate sector, when handling foreign-related business, translation functions are usually required. Existing text translation methods usually adopt an end-to-end direct translation mode, that is, the input text is processed as a whole through a general translation model or software. However, since the translation model or software lacks knowledge base support for the real estate field, it is difficult to accurately translate information texts related to the real estate field. In addition, information texts related to the real estate field usually contain mixed content of multiple fields. If the entire text is translated directly, it is easy to cause professional fields in the real estate field to be mixed with general content, resulting in mistranslation of terms or loss of format, reducing the accuracy of the translation and affecting the user experience. Therefore, the current text translation for the real estate field has the problem of poor translation accuracy.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a text translation method, system, electronic device, storage medium and computer program product, aiming to solve the technical problem of poor translation accuracy.
[0005] To achieve the above objectives, the present application proposes a text translation method, which includes:
[0006] Identify the original property information text in the property information image;
[0007] Segmenting the original property information text to obtain text fields, and translating the text fields into a target language using a first preset large model to obtain a first translation field;
[0008] Using a second preset large model, the first translation field is detected to determine whether it conforms to the usage scenario of the target language, thereby obtaining translation guidance, and the text field is modified based on the translation guidance to obtain a second translation field.
[0009] The target property information text is obtained according to the second translation field.
[0010] In one embodiment, the step of identifying the original property information text in the property information image includes:
[0011] Detecting a text block in the house listing image and determining a layout rule of the text block;
[0012] The text content in the text block is extracted by optical character recognition technology, and the logical order of the text content is adjusted in combination with the typesetting rules to obtain the original house information text.
[0013] In one embodiment, the target language usage scenario includes a katakana translation index, a sentence structure translation index, and a unit translation index, the translation guidance includes a first translation guidance, and the step of detecting whether the first translation field conforms to the target language usage scenario using a second preset large model, and obtaining the translation guidance includes:
[0014] If the second preset large model identifies that the translation of the katakana in the text field by the first translation field is not a transliteration or a corresponding translated name in the target language, determining that the first translation field does not meet the katakana translation index, and generating a katakana translation difference for the first translation field;
[0015] Identifying the sentence structure of the first translation field using the second preset large model, and if a component of the sentence structure is missing and affects the meaning of the first translation field, determining that the first translation field does not meet the sentence structure translation indicator, and generating a sentence structure translation difference for the first translation field;
[0016] If the second preset large model identifies that the translation of the unit in the text field by the first translation field does not conform to the corresponding unit in the target language, determining that the first translation field does not conform to the unit translation indicator, and generating a unit translation difference for the first translation field;
[0017] A first translation guidance opinion is obtained by combining the katakana translation difference, the sentence structure translation difference and the unit translation difference.
[0018] In one embodiment, the target language usage scenario further includes an expression translation index, a semantic translation index, and a logical translation index, the translation guidance further includes a second translation guidance, and the step of detecting whether the first translation field conforms to the target language usage scenario using the second preset large model and obtaining the translation guidance further includes:
[0019] If the second preset large model identifies that the expression of the first translation field does not conform to the polite expression of the target language, determining that the first translation field does not conform to the expression translation indicator, and generating an expression translation difference for the first translation field;
[0020] When the second preset large model identifies whether the translation of the idiom or phrase in the text field by the first translation field conforms to the corresponding semantic expression in the target language, determining that the first translation field does not conform to the semantic translation indicator, and generating a semantic translation difference for the first translation field;
[0021] When the second preset large model identifies that the word order of the first translation field is disordered or the meaning of the sentence is unclear, determining that the first translation field does not meet the logical translation indicator, and generating a logical translation difference for the first translation field;
[0022] A second translation guidance opinion is obtained by combining the expression translation difference, the semantic translation difference and the logical translation difference.
[0023] In one embodiment, the target language usage scenario further includes a proper noun translation indicator, a housing information translation indicator, and a unit structure translation indicator; the translation guidance further includes a third translation guidance; and the step of detecting whether the first translation field conforms to the target language usage scenario using the second preset large model and obtaining the translation guidance further includes:
[0024] If the second preset large model identifies that the translation of the proper noun in the text field by the first translation field does not conform to the standard expression of proper nouns in the real estate field, determining that the first translation field does not conform to the proper noun translation index, and generating a proper noun translation difference for the first translation field;
[0025] If the second preset large model identifies that the translation of the property information in the text field by the first translation field does not conform to the standard expression of information in the real estate field, determining that the first translation field does not conform to the property information translation index, and generating a property information translation difference for the first translation field;
[0026] When the second preset large model identifies that the translation of the apartment structure in the text field by the first translation field does not conform to the standard expression of apartment structure in the real estate field, determining that the first translation field does not conform to the apartment structure translation index, and generating an apartment structure translation difference for the first translation field;
[0027] A third translation guidance opinion is obtained by combining the translation differences of the proper nouns, the translation differences of the housing information, and the translation differences of the apartment structure.
[0028] In one embodiment, the step of obtaining the target property information text according to the second translation field includes:
[0029] Integrating the second translation field into a translation text;
[0030] The corresponding portion of the original property information text in the translated text is adjusted according to a preset field dictionary.
[0031] In one embodiment, the step of adjusting the corresponding portion of the original property information text in the translated text according to a preset field dictionary includes:
[0032] Extracting system fields from the original property information text;
[0033] A match is performed in a preset field dictionary based on the system field to obtain a standard translation field corresponding to the system field, and a portion corresponding to the system field in the translation text is replaced with the standard translation field.
[0034] In addition, to achieve the above objectives, the present application also proposes a text translation system, which includes:
[0035] A text recognition module is used to identify the original property information text in the property information image;
[0036] A first translation module is configured to segment the original property information text to obtain text fields, and translate the text fields into a target language using a first preset large model to obtain a first translation field;
[0037] A second translation module is configured to detect whether the first translation field conforms to the usage scenario of the target language using a second preset large model, obtain translation guidance, and modify the text field based on the translation guidance to obtain a second translation field;
[0038] A field adjustment module is used to obtain the target property information text according to the second translation field.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text translation method described above.
[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the steps of the text translation method described above are implemented.
[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the text translation method described above are implemented.
[0042] The present application provides a text translation method, which includes: identifying original property information text in a property information picture; segmenting the original property information text to obtain text fields, and translating the text fields into a target language using a first preset large model to obtain a first translation field; detecting whether the first translation field conforms to the usage scenario of the target language using a second preset large model, obtaining translation guidance, and correcting the text field based on the translation guidance to obtain a second translation field; and obtaining the target property information text based on the second translation field.
[0043] This application realizes the conversion from image to text by receiving a property information picture and identifying the original property information text therein, providing basic data for subsequent text translation. By segmenting the original property information text to obtain text fields, and then translating these fields separately to obtain the first translation field, the problems caused by translating the entire text are avoided, and the accuracy of the translation is initially improved. By detecting whether the first translation field meets the translation evaluation criteria, the problems therein can be discovered and translation guidance opinions can be obtained. Based on these guidance opinions, the text field is translated for the second time to obtain the second translation field, further improving the accuracy of the translation. The second translation field is integrated into the translation text, and the corresponding part of the original property information text in the translation text is adjusted according to the preset field dictionary. This not only ensures that the format and structure of the translated text are consistent with the original property information, improving the readability and usability of the text, but also ensures that the content of the target property information text complies with the provisions of the preset field dictionary, thereby improving the accuracy of the translation. Compared with existing text translation methods that usually use an end-to-end direct translation mode to process the input text as a whole, which easily leads to the mixing of professional fields and general content, and lacks knowledge base support for the real estate field, this solution can better handle professional fields in the real estate field through segmented processing combined with secondary translation, making the translation results more in line with the professional requirements of the real estate field, enhancing the pertinence and professionalism of the translation, providing users with more accurate and reliable housing information translation results, improving user experience, and promoting the smooth progress of foreign-related real estate business. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1A flowchart of the first embodiment of the text translation method of this application is provided;
[0047] Figure 2 A flowchart of the second embodiment of the text translation method of this application is provided;
[0048] Figure 3 An overall flow chart of the method for translating the application text;
[0049] Figure 4 This is a schematic diagram of the module structure of the text translation system according to an embodiment of the present application;
[0050] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the text translation method in the embodiment of the present application.
[0051] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0053] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0054] The main solution of the embodiment of the present application is: identifying the original property information text in the property information picture; segmenting the original property information text to obtain text fields, and translating the text fields into the target language through a first preset large model to obtain a first translation field; detecting whether the first translation field conforms to the usage scenario of the target language through a second preset large model, obtaining translation guidance, and correcting the text field based on the translation guidance to obtain a second translation field; and obtaining the target property information text based on the second translation field.
[0055] In this embodiment, for ease of description, the text translation system is used as the execution subject for explanation.
[0056] Since the existing technology usually adopts an end-to-end direct translation mode to process the input text as a whole, it is easy to cause professional fields and general content to be mixed together, and there is a lack of knowledge base support for the real estate field. This application provides a solution that, through segmented processing combined with secondary translation, can better handle professional fields in the real estate field, make the translation results more in line with the professional requirements of the real estate field, enhance the pertinence and professionalism of the translation, provide users with more accurate and reliable housing information translation results, improve user experience, and promote the smooth progress of foreign-related real estate business.
[0057] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or a text translation system, application app, mini-program, etc. that can implement the above functions. The following text translation system is used as an example to illustrate this embodiment and the following embodiments.
[0058] Based on this, the present application embodiment provides a text translation method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the text translation method of this application.
[0059] In this embodiment, the text translation method includes steps S01 to S04:
[0060] Step S01, identifying the original property information text in the property information image;
[0061] It should be noted that the original property information text is extracted from the property information image through optical character recognition (OCR) technology. The property information image is an image file containing property information, such as a house advertisement, a scanned copy of a real estate certificate, etc. The property information may include the house address, rent, apartment type, etc. The original property information text is the unprocessed text content extracted from the image, and is the text form of the property information.
[0062] It is understandable that in foreign-related business in the real estate sector, property information often exists in the form of pictures (such as real estate advertisements, scanned copies of real estate certificates, etc.), and existing translation methods only support plain text input and cannot directly process picture content. Therefore, step S01 is performed to convert the property information in the picture into editable original text to provide basic data for subsequent translation.
[0063] Step S02: Segment the original property information text to obtain text fields, and translate the text fields into a target language using a first preset large model to obtain a first translation field;
[0064] It should be noted that the field boundaries of the original property information text are identified, where the field boundaries can be identified based on punctuation marks (for example, commas, semicolons, periods, etc. as separators), keywords (for example, real estate keywords such as "area", "price", and "unit type"), and template matching (for example, the predefined format of "address-unit type-area-price"). The identified fields are cleaned and standardized, redundant information is removed, and the units and formats are unified to obtain independent units in the original property information text that are divided by semantics or format, namely text fields.
[0065] In addition, it should be noted that to obtain text fields, it is also possible to organize the scattered original property information text into structured text fields by calling a large model such as GPT4o or deepseek. After obtaining the text fields, a first preset large model is used to translate the language corresponding to each original text field into the target language, for example, translating Japanese into Chinese. The result of the initial translation is the first translated field. The first preset large model is a generative large model such as GPT4o or deepseek.
[0066] It is understandable that since the existing method directly translates the entire text, it is easy to cause professional terms and general content to be mixed, resulting in terminology mistranslation. Therefore, step S02 is performed to separate professional terms and general content through field segmentation and translate them separately, reducing the translation model's dependence on context, thereby reducing the risk of terminology mistranslation.
[0067] Step S03: Using a second preset large model to detect whether the first translation field conforms to the target language usage scenario, obtaining translation guidance, and modifying the text field based on the translation guidance to obtain a second translation field;
[0068] It should be noted that the standardization of the translation of the first translation field is tested using a second preset large model to determine whether it conforms to the target language usage scenario. The second preset large model is a generative large model and can be the same as or different from the first preset large model. The target language usage scenario is a series of indicators used to measure translation quality, used to evaluate the standardization of translation in terms of grammar, cultural differences, and the real estate sector. In terms of grammar, the first translation field is primarily evaluated for grammatical norms, syntactic structure, and unit symbol usage. In terms of cultural differences, the first translation field is primarily evaluated for polite expressions and the expression of idioms and common expressions. In the real estate sector, the first translation field is primarily evaluated for the expression of proper place names or nouns, housing information, and unit type information. By testing whether the first translation field conforms to the target language usage scenario, non-standard areas of the first translation field can be identified, and translation guidance is generated based on these non-standard areas. Keyword extraction is performed on the translation guidance, and the extracted keywords are input as prompt words into the second preset large model as constraints for secondary correction. The text field is then retranslated to obtain the secondary translation result, namely the second translation field.
[0069] It is understandable that since the existing solution relies on a general translation model and lacks support from a domain knowledge base, step S03 is performed to detect whether the first translation field conforms to the usage scenario of the target language, evaluate the accuracy of the first translation field, and generate translation guidance. Based on the translation guidance, the non-standard translation parts are corrected to further improve the translation quality and achieve closed-loop optimization of translation quality.
[0070] Step S04: Obtain target property information text according to the second translation field.
[0071] It should be noted that after obtaining all the second translation fields, the second translation fields are connected into a complete translation text in the field sorting order. The translation text corresponds to the original property information text. The part of the original property information text corresponding to the preset field dictionary is determined, and the translation text of this part is adjusted to make it conform to the content in the preset field dictionary. The preset field dictionary contains mapping rules for specific terms and formats in the real estate field. The target property information text is the final translation result after adjustment.
[0072] It is understandable that since the large model may generate information that seems reasonable but does not conform to facts, logic or contextual semantics when outputting content (i.e., model hallucination), step S04 is performed. By matching the preset field dictionary and executing the field replacement step, it can ensure that the expression logic of the target property information text is self-consistent, avoid cross-paragraph contradictions, and conform to the common expressions in the real estate field, thereby achieving the effect of suppressing model hallucinations.
[0073] In a feasible implementation, in step S01, the step of identifying the original property information text in the property information image includes steps A01 to A02:
[0074] Step A01: Detect text blocks in the house listing image and determine the layout rules of the text blocks;
[0075] It should be noted that deep learning-based object detection algorithms (such as YOLO or Faster R-CNN) or large models are used to perform pixel-level analysis of property images to locate contiguous regions containing text (i.e., text blocks). These text blocks may correspond to different pieces of information in the property image, such as the title, address, price, and description. The relative position (e.g., coordinates, spacing, and alignment) and semantic associations (e.g., "price" and "area" are often adjacent) of the text blocks are analyzed to determine the layout rules for the text blocks, including horizontal and vertical layouts.
[0076] For example, if the text blocks are arranged vertically and the spacing is even, it is judged to be vertical layout (such as some old real estate certificates); if there is a title row (such as "Property Information Table") and the content below is aligned in columns, it is judged to be table layout; if multiple text blocks are evenly arranged in the vertical direction and the font size is the same, they belong to the same column of information; if the text blocks are horizontally aligned and the spacing is small, they belong to the same row of information.
[0077] Step A02: extract the text content in the text block using optical character recognition technology, and adjust the logical order of the text content in combination with typesetting rules to obtain the original property information text.
[0078] It should be noted that OCR technology is used to convert the pixel information in the text block into a character string to obtain the text content in the text block, and the extracted text content is re-arranged in a logical order (i.e., the semantically correct arrangement order of the text content) in combination with typesetting rules.
[0079] For example, if the typesetting rule is "tabular typesetting", the text content is reorganized in row and column order; if the typesetting rule is "vertical typesetting", the text blocks are spliced in order from top to bottom and from left to right; if the typesetting rule is "horizontal typesetting", the text blocks are spliced in order from left to right.
[0080] In this embodiment, by accurately identifying the text area in the image and filtering out non-text elements to ensure the purity of subsequent processing, the typesetting rules of the text block are inferred by using image features and semantic analysis, providing a structured basis for subsequent logical reorganization. By predicting the typesetting rules, text blocks with abnormal typesetting (such as blurred or severely tilted areas) can be filtered out in advance, thereby improving the robustness of text extraction. Based on OCR technology, complex fonts, handwriting and low-resolution images are recognized, significantly improving the accuracy of text extraction. Combined with typesetting rules, the fragmented text extracted by OCR is re-arranged according to structures such as titles, paragraphs, and tables to generate logically coherent original text. The text language is automatically identified through the language detection model, and the corresponding OCR module is called to achieve seamless processing of multilingual text.
[0081] In a feasible implementation, in step S03, the target language usage scenario includes a katakana translation index, a sentence structure translation index, and a unit translation index; the translation guidance includes a first translation guidance; and the second preset large model is used to detect whether the first translation field meets the target language usage scenario. The step of obtaining the translation guidance includes steps A11 to A14:
[0082] Step A11: when the second preset large model identifies that the translation of the katakana in the text field by the first translation field is not a transliteration or a corresponding translated name in the target language, determining that the first translation field does not meet the katakana translation index, and generating a katakana translation difference for the first translation field;
[0083] It should be noted that the katakana part in the first translation field is extracted through regular expressions or tokenization tools. Katakana is often used for transliteration of foreign words or industry names or translation of specific translated names. The katakana part and the limit of the katakana translation index are input into the second preset large model. The katakana translation index is used to measure the accuracy of the katakana translation in the first translation field for the text field. It is required to translate according to transliteration or the corresponding translated name in the target language. Based on the second preset large model, it is detected whether the transliteration rule of the katakana part conforms to the target phonetic system. When detecting, the preset transliteration rule library (which contains common katakana and corresponding translations) can be called to forcibly match the translation name with the highest priority. For the katakana not registered in the preset transliteration rule library in the text field, calculate the BERT embedding cosine similarity between it and other translated names in the preset transliteration rule library. The BERT embedding cosine similarity measures the similarity between two vectors (usually the embedding vectors of words or sentences) by calculating the included angle between them, determine the translated name with the highest BERT embedding cosine similarity, and use it as the translated name with the highest priority. Determine whether the translated name with the highest priority is consistent with the translation of the first translation field. In the case of inconsistency, it is determined that the first translation field does not meet the katakana translation index, and a katakana translation difference is generated, which explains the difference between the katakana translation in the first translation field and the correct translation and the correction suggestions. For example, "コーヒー" is wrongly translated as "可非", but it should actually be translated as "咖啡".
[0084] Exemplarily, when the language corresponding to the original text field is Japanese and the target language is Chinese, the katakana in the text field is "エレベーター", and the correct translation is "电梯". If the first translation field translates it as "电升梯", there is a katakana translation difference, that is, the difference between "电升梯" and "电梯".
[0085] Step A12, identify the sentence structure of the first translation field through the second preset large model. In the case where a component of the sentence structure is missing and affects the expression of the meaning of the first translation field, it is determined that the first translation field does not meet the sentence structure translation index, and a sentence structure translation difference of the first translation field is generated;
[0086] It should be noted that the syntactic tree structure of the first translation field is parsed through natural language processing (NLP) technology, and the restrictions of sentence structure translation metrics are input into the second preset large model. The sentence structure translation metrics are used to evaluate whether the sentence structure of the first translation field is complete and reasonable. Based on the second preset large model, constituent analysis of the first translation field is performed to determine whether the sentence constituents are complete (whether there are a subject, a predicate, and an object), and the lexical collocations are verified through a preset semantic knowledge base (for example, the verb "eat" is usually followed by food), so as to determine whether the meaning expression logic of the first translation field is reasonable. When making the judgment, a pre-trained machine learning model can also be combined for judgment. If the sentence constituents of the first translation field are incomplete and the meaning expression logic is unreasonable, it is determined that the first translation field does not meet the sentence structure translation metrics, and a sentence structure translation difference is generated, which explains the problems existing in the sentence structure of the first translation field (such as missing components, unreasonable structure, etc.) and the correction suggestions.
[0087] Exemplarily, when the language of the original text field is Japanese and the target language is Chinese, the Japanese sentence structure of the text field is "私はりんごを食べます" (I eat apples, subject-predicate-object structure). If the first translation field is "苹果吃我", the sentence structure translation difference is obvious and does not conform to the correct subject-predicate-object structure of Chinese.
[0088] Step A13, when the second preset large model identifies that the translation of the unit in the first translation field for the text field does not conform to the unit corresponding to the target language, it is determined that the first translation field does not meet the unit translation metrics, and a unit translation difference of the first translation field is generated;
[0089] It should be noted that for the unit in the boundary text field between numerical characters and non-alphanumeric characters (such as spaces, slashes), for example, for an area of 60.5㎡, the extraction result is 60.5 (numerical value) + square meters (unit). The restrictions of the unit translation metrics are input into the second preset large model. The unit translation metrics are used to check whether the translation of the unit in the first translation field conforms to the unit corresponding to the target language. Based on the second preset large model, it is checked whether the translation of the unit in the first translation field is consistent with the translation of the unit corresponding in the preset unit conversion rule library. In the case of inconsistency, it is determined that the first translation field does not meet the unit translation metrics, and a unit translation difference is generated, which explains the wrongly translated unit in the first translation field and the correct unit of the target language.
[0090] Exemplarily, when the language corresponding to the original text field is Japanese and the target language is Chinese, the unit in the text field is "円" (Japanese yen), which should be correctly translated as "Japanese yen". If the first translation field translates it as "圆" (although "圆" can also mean "Japanese yen" in some informal situations, but according to strict language grammar standards, the standard translation is "Japanese yen"), there is a unit translation difference.
[0091] Step A14, combine the katakana translation difference, sentence structure translation difference and unit translation difference to obtain the first translation guidance.
[0092] It should be noted that the katakana translation difference, sentence structure translation difference and unit translation difference information obtained in steps A11 to A13 are integrated to generate a comprehensive first translation guidance, which points out the problems existing in the first translation field under the language grammar standards and the corresponding correction suggestions.
[0093] In this embodiment, by detecting the katakana translation difference, the translation accuracy of special vocabulary is significantly improved, and the semantic ambiguity caused by incorrect phonetic correspondence is reduced. By parsing the syntactic tree structure, problems such as word order deviation and component omission can be identified and corrected to ensure that the translated text conforms to the grammar norms of the target language and improve the accuracy at the sentence level. By identifying the unit translation deviation, the semantic consistency between the numerical value and the unit is ensured, and the semantic ambiguity caused by incorrect units is avoided. By integrating the analysis results of the differences in katakana, sentence structure and unit translation, a comprehensive guidance is generated, significantly reducing the system correction cost and improving the translation efficiency and accuracy.
[0094] In a feasible embodiment, in step S03, the usage scenarios of the target language also include expression translation indicators, semantic translation indicators, and logical translation indicators. The translation guidance also includes the second translation guidance. The steps of obtaining the translation guidance by detecting whether the first translation field conforms to the usage scenario of the target language by the second preset large model also include steps A21 to A24:
[0095] Step A21, when the second preset large model identifies that the expression of the first translation field does not conform to the polite expression of the target language, it is determined that the first translation field does not conform to the expression translation indicator, and the expression translation difference of the first translation field is generated;
[0096] It should be noted that the politeness level of the first translation field is detected and classified by the politeness level analyzer of the second preset large model, and the first translation field is analyzed based on the politeness expression norms of the target language (such as honorific systems, euphemism usage rules, etc.). When it is detected that the politeness level of the first translation field does not meet the standard (for example, the politeness level is "impolite") or does not conform to the politeness expression norms of the target language (such as too direct tone, lack of honorifics, etc.), it is determined that the first translation field does not meet the expression mode translation index. The expression mode translation index is used to measure whether the translation result conforms to the expression habits of the target language (such as politeness level, tone, etc.). At the same time, an expression mode translation difference is generated, which explains the honorifics to be added or the tone adjustment strategy in the first translation field.
[0097] Exemplarily, when the language corresponding to the original text field is Japanese and the target language is Chinese, "お客様、お願いします。ドアを閉めてください。" is directly translated as "Customer, please. Close the door." Although this translation is basically correct semantically, from the perspectives of tone and cultural habits, compared with "Hello, could you please help me close the door, thank you!", which is more in line with the reading habits of the target language, there is a language expression translation difference that does not conform to the Chinese politeness expression habits.
[0098] Step A22, when the second preset large model identifies whether the translation of idiomatic expressions or idioms in the first translation field for the text field conforms to the corresponding semantic expression of the target language, it is determined that the first translation field does not meet the semantic translation index, and a semantic translation difference of the first translation field is generated;
[0099] It should be noted that by comparing the semantic similarity between the first translation field and idiomatic expressions or idioms through the second preset large model, or by matching the idiom / metaphor corresponding translations in the preset semantic image correspondence table, when the semantic similarity does not reach the preset similarity threshold (for example, 90%) or the translation of idiomatic expressions or idioms in the first translation field for the text field has no matching translation in the preset semantic image correspondence table, it is determined that the first translation field does not meet the semantic translation index. The semantic translation index is used to measure whether the translation result accurately conveys the semantics of the original text. At the same time, a semantic translation difference is generated, which records the inaccurate part and the correct expression of the idiomatic expressions or idioms in the first translation field.
[0100] Exemplarily, when the language corresponding to the original text field is Japanese and the target language is Chinese, translating "猫の手も借りたい" into other inappropriate expressions results in the inability to accurately convey the semantic image of the original sentence "extremely busy".
[0101] Step A23: When the second preset large model identifies that the word order of the first translation field is disordered or the sentence meaning is unclear, it determines that the first translation field does not meet the logical translation criteria and generates the logical translation difference of the first translation field.
[0102] It should be noted that the deviation degree of the first translation field from the normal word order is calculated by the second preset large model. When the deviation degree is higher than the preset deviation threshold, it is determined that the word order of the first translation field is disordered. At the same time, the second preset large model is also used to identify whether there is contradictory information in the first translation field. When there is contradictory information, it is determined that the sentence meaning of the first translation field is unclear. When the first translation field has problems such as disordered word order or contradictory information, it is determined that the first translation field does not meet the logical translation criteria, and at the same time, the logical translation difference is generated, including the expression problems and correct expressions that exist logically in the first translation field.
[0103] Exemplarily, when the original text field corresponds to Japanese and the target language is Chinese, "どちらがいいか、考えています" is directly translated as "正在想哪个更好" (which is "thinking about which one is better" in Chinese). In the Chinese context, this sentence structure will seem incomplete. A more common translation is "我正在想哪个选择更合适" (which is "I'm thinking about which option is more suitable").
[0104] Step A24: Combine the expression translation difference, semantic translation difference, and logical translation difference to obtain the second translation guidance.
[0105] It should be noted that the language expression translation difference, semantic image translation difference, and sentence structure translation difference obtained in steps A21 to A23 are integrated to generate a comprehensive second translation guidance. This guidance points out the problems existing in the first translation field under the cultural difference standard and the corresponding correction suggestions.
[0106] In this embodiment, by determining the expression translation difference (such as identifying a more natural expression in the target language), the guidance system adjusts the diction to make the translation more in line with the daily expression habits of the target language, thereby improving the naturalness and readability of the translation. By identifying the semantic translation difference, the guidance system adopts equivalent cultural images or supplementary explanatory notes in the target language to ensure that the original meaning is retained in the target culture, avoiding ambiguity or cultural misunderstandings. By adjusting the logical translation difference, the translation conforms to the grammar habits and logical order of the target language, improving the fluency and logic of the translation, and reducing the understanding obstacles caused by improper sentence patterns. By integrating the three types of differences, a multi-dimensional translation guidance is generated, providing an accurate correction direction for the system, significantly improving the translation accuracy, and achieving semantic, pragmatic, and discourse-level equivalence of the translation in the target culture.
[0107] In a feasible implementation manner, in step S03, the usage scenarios of the target language also include the proper noun translation index, the housing source information translation index, and the housing type structure translation index. The translation guidance also includes the third translation guidance. The step of detecting whether the first translation field conforms to the usage scenario of the target language through the second preset large model further includes steps A31 to A34:
[0108] Step A31, in the case where the second preset large model identifies that the translation of the proper noun in the text field by the first translation field does not conform to the standard expression of the proper noun in the real estate field, it is determined that the first translation field does not conform to the proper noun translation index, and a proper noun translation difference of the first translation field is generated;
[0109] It should be noted that in the real estate field, the translation of proper nouns needs to follow specific norms and usage habits to ensure the accuracy and consistency of information transmission. The proper noun translation index covers multiple aspects such as place names, building names, and professional terms, and is an important basis for translation. Proper nouns refer to nouns with specific meanings and referents in the real estate field, such as specific building names, place names, professional terms, etc. For example, "渋谷ヒルズ" is a specific building name referring to "Shibuya New City", and "マンション" is a proper noun in Japanese representing an apartment. By adding domain restrictions of the real estate field to the second preset large model, it is judged whether the translation of the proper noun in the text field by the first translation field conforms to the expression under the real estate field standard. It can also be judged according to a preset proper noun correspondence table, which contains common proper nouns and corresponding translations. If the translation of the proper noun in the text field by the first translation field does not conform to the standard expression under the real estate field standard, it is determined that the first translation field does not conform to the proper noun translation index, and a proper noun translation difference is generated, including the proper noun and the corresponding translation.
[0110] Exemplarily, when the language corresponding to the original text field is Japanese and the target language is Chinese, "渋谷ヒルズ" should be translated as "Shibuya New City", "マンション" as "apartment", and "一戸建て" as "single-family house". When the first translation field wrongly translates "渋谷ヒルズ" as "Shibuya Building", through matching detection with the preset proper noun correspondence table, it can be identified that the translation of this proper noun does not conform to the real estate field standard, thus determining the proper noun translation difference.
[0111] Step A32, in the case where the second preset large model identifies that the translation of the housing source information in the text field by the first translation field does not conform to the information standard expression in the real estate field, it is determined that the first translation field does not conform to the housing source information translation index, and a housing source information translation difference of the first translation field is generated;
[0112] It should be noted that property information refers to various specific details about the property itself within a real estate project and is an important basis for users to understand the basic details of the property, such as the property's transportation and facility details, and the property's square footage. Based on the standard translation requirements for property information in the real estate sector, the property information in the first translation field is tested using a second preset large model to determine whether the first translation field's translation of the text field is standard. The translation is also compared with a standard property information translation template to determine the accuracy of the translation. If the first translation field's translation of the property information in the text field does not meet the standard translation requirements for property information in the real estate sector, the first translation field is determined to not meet the property information translation indicators, and a property information translation difference is generated, including the property information and corresponding translation details.
[0113] For example, property information typically includes area, price, and ownership period. If the first translation field incorrectly translates the area unit "square meters" as "square feet," or the price figure is incorrectly translated, a machine learning model can identify the non-compliant portion of the property information and identify the translation discrepancy, such as "The area unit square meters should be accurately translated, but was incorrectly translated as square feet." Furthermore, the translation process must ensure that all relevant details, such as proximity to subways, schools, and shopping malls, are preserved and accounted for.
[0114] Step A33: When the second preset large model identifies that the translation of the apartment structure in the text field by the first translation field does not conform to the standard expression of apartment structure in the real estate field, it is determined that the first translation field does not conform to the apartment structure translation index, and an apartment structure translation difference is generated for the first translation field;
[0115] It should be noted that "household structure" refers to the internal layout and functional zoning of a house and is an important concept for describing the spatial characteristics of a house. Based on the translation rules for "household structure" in real estate standards, the house structure in the first translation field is tested using a second preset large model. Alternatively, the test can be performed based on pre-set house structure translation rules. If the first translation field's translation of the "household structure" in the text field does not conform to the translation rules for "household structure" in real estate standards, the first translation field is determined to be non-compliant with the "household structure" translation indicator, and a "household structure" translation difference is generated, including the "household structure" and its corresponding standardized expression.
[0116] For example, when the language corresponding to the original text field is Japanese and the target language is Chinese, when the apartment structure is units such as "LDK" and "DK", no translation is performed according to the standard; if the first translation field mistakenly translates it into other content, or the translation of other common apartment structures such as "three bedrooms and one living room" is inaccurate, by matching with the pre-set apartment structure translation rules, it can be determined that the translation of the apartment structure does not meet the standards in the real estate field, and the apartment structure translation difference is obtained.
[0117] In step A34, third translation guidance is obtained by combining the translation differences of proper nouns, the translation differences of housing information, and the translation differences of apartment structures.
[0118] It should be noted that the noun translation differences, property information translation differences, and apartment structure translation differences obtained in steps A31 to A33 are integrated to generate a comprehensive third translation guidance opinion, which points out the problems existing in the first translation field under the real estate field standards and corresponding correction suggestions.
[0119] In this embodiment, by clearly identifying the differences in proper noun translation, the problems existing in the proper noun translation of the first translation field can be clearly demonstrated, providing a clear direction for subsequent corrections, helping to improve the accuracy of proper noun translation and ensure the correct transmission of professional information. By accurately identifying the differences in property information translation, it helps to ensure that property information is not distorted during the translation process, enabling the target audience to accurately understand the key features of the property, improving the accuracy and effectiveness of property information dissemination, and providing reliable information support for real estate transactions. By clarifying the differences in apartment structure translation, the translation results can more accurately convey the apartment information of the house, helping users better understand the house layout, and also providing a more accurate basis for information exchange in the real estate market. By combining the differences in proper noun translation, property information translation, and apartment structure translation, a comprehensive analysis of translation issues in multiple aspects can be conducted, providing a more comprehensive and specific direction for improvement, helping the system to grasp the overall translation quality, make targeted corrections, improve the accuracy and professionalism of translation, effectively solve the problem of poor translation accuracy, and improve the overall quality of text translation in the real estate field.
[0120] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 In step S04, the step of obtaining the target property information text according to the second translation field includes steps S11 to S12:
[0121] Step S11, integrating the second translation field into a translation text;
[0122] It should be noted that the second translation field is reorganized, conjunctions and punctuation are inserted to ensure semantic coherence, and the paragraph order is restored according to the original text typesetting logic to obtain a complete translation text, which contains all the property information and conforms to grammar and logic.
[0123] Step S12: adjusting the corresponding portion of the original property information text in the translated text according to the preset field dictionary to obtain the target property information text.
[0124] It should be noted that based on the special term library (preset field dictionary) in the real estate field, mandatory replacement and format calibration are performed on the non-fully standardized parts in the translated text to generate the final target housing source information text.
[0125] In this embodiment, by integrating the translation fragments, a structured translation text is generated to ensure that the translation content of each field is complete and unified in style. Through the preset field dictionary, the terms in the translation text are unified into standard terms, reducing term errors, making it meet the business requirements, reducing ambiguity and errors in translation, and significantly improving the accuracy of translation.
[0126] In a feasible embodiment, in step S04, the steps of adjusting the corresponding part of the original housing source information text in the translation text according to the preset field dictionary include steps B11 to B12:
[0127] Step B11, extract the system fields in the original housing source information text;
[0128] It should be noted that the system fields in the original text are identified and extracted through machine learning algorithms or large models. System fields refer to key information fields with fixed formats or specific identifiers in the original housing source information text, such as "House Name", "Basic Information", etc. These fields need to maintain their accuracy and consistency during the translation process.
[0129] Step B12, based on the system fields, perform matching in the preset field dictionary to obtain the standard translation fields corresponding to the system fields, and replace the corresponding part of the system fields in the translation text with the standard translation fields.
[0130] It should be noted that these system fields are matched with the corresponding standard translation fields in the preset field dictionary. Standard translation fields refer to the translated versions corresponding to the system fields in the preset field dictionary. These fields are common terms in the real estate field and are mandatory and unchanged, and their meanings will not change with model translation. After the matching is completed, the corresponding part of the system fields in the translation text is replaced with these standard translation fields, ensuring the accuracy and standardization of the housing source information in the translation text, and improving the efficiency and accuracy of information processing.
[0131] Exemplarily, when the language of the original text field is Japanese and the target language is Chinese, the system fields in the original housing source information text are "建物の名前", "建物の名前です", "物件名", "建物名", and the corresponding standard translation fields in the preset field dictionary are all "House Name".
[0132] In this embodiment, by extracting the system fields in the original property information text, a solid foundation is laid for the correct identification and replacement of system fields in the subsequent translation process. It can also greatly improve the accuracy of translation and reduce translation problems caused by field recognition errors. By introducing a preset field dictionary, the system fields are matched with the corresponding standard translation fields in the dictionary, and the corresponding parts of the system fields in the translated text are replaced with standard translation fields, ensuring that the translation of the system fields is accurate and greatly improving the quality and accuracy of the translated text. At the same time, the use of a preset field dictionary can also improve the consistency of translation and avoid the problem of inconsistent translation expressions of the same system fields in different translation tasks. In addition, accurate translation text can better convey information, improve the convenience of user reading and understanding, and thus optimize the user experience.
[0133] For example, to help understand the technical concept or technical principle of this application, please refer to Figure 3 , Figure 3 A comprehensive flowchart of the translation method is provided. After uploading a property information image, the text is extracted using OCR technology to obtain the original property information text. The property information text is then segmented. The segmented text fields are initially translated to obtain the first translated field. The translation result (i.e., the first translated field) is then corrected to determine whether it meets the translation evaluation criteria. Based on the translation guidance obtained from the test, the text field is then translated again to obtain the second translated field. Field dictionary matching is then used to adjust the corresponding portion of the original property information text in the translated text, ultimately outputting the target property information text.
[0134] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the text translation method of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.
[0135] This application also provides a text translation system, please refer to Figure 4 , the text translation system includes:
[0136] A text recognition module 10 is used to recognize the original property information text in the property information image;
[0137] A first translation module 20 is configured to segment the original property information text to obtain text fields, and translate the text fields into a target language using a first preset large model to obtain a first translation field;
[0138] A second translation module 30 is configured to detect whether the first translation field conforms to the target language usage scenario using a second preset large model, obtain translation guidance, and modify the text field based on the translation guidance to obtain a second translation field;
[0139] The field adjustment module 40 is used to obtain the target property information text according to the second translation field.
[0140] Optionally, the text recognition module 10 is further configured to:
[0141] Detect text blocks in property information images and determine the layout rules of the text blocks;
[0142] The text content in the text block is extracted through optical character recognition technology, and the logical order of the text content is adjusted in combination with typesetting rules to obtain the original property information text.
[0143] Optionally, the target language usage scenario includes a katakana translation index, a sentence structure translation index, and a unit translation index, the translation guidance includes a first translation guidance, and the second translation module 30 is further configured to:
[0144] When the second preset large model identifies that the translation of the katakana in the text field by the first translation field is not a transliteration or a corresponding translated name in the target language, determining that the first translation field does not meet the katakana translation index, and generating a katakana translation difference for the first translation field;
[0145] Identify the sentence structure of the first translation field using the second preset large model, and if a component of the sentence structure is missing and affects the meaning of the first translation field, determine that the first translation field does not meet the sentence structure translation index, and generate a sentence structure translation difference for the first translation field;
[0146] When the second preset large model identifies that the translation of the unit in the text field by the first translation field does not conform to the corresponding unit in the target language, determining that the first translation field does not conform to the unit translation indicator, and generating a unit translation difference for the first translation field;
[0147] The first translation guidance is obtained by combining the Katakana translation differences, sentence structure translation differences and unit translation differences.
[0148] Optionally, the target language usage scenario further includes an expression translation indicator, a semantic translation indicator, and a logical translation indicator, and the translation guidance further includes a second translation guidance. The second translation module 30 is further configured to:
[0149] When the second preset large model identifies that the expression of the first translation field does not conform to the polite expression of the target language, it is determined that the first translation field does not conform to the expression translation index, and an expression translation difference of the first translation field is generated;
[0150] When the second preset large model identifies whether the translation of the idiom or phrase in the text field by the first translation field conforms to the corresponding semantic expression in the target language, determining that the first translation field does not conform to the semantic translation index, and generating a semantic translation difference for the first translation field;
[0151] When the second preset large model identifies that the word order of the first translation field is disordered or the meaning of the sentence is unclear, it is determined that the first translation field does not meet the logical translation index, and a logical translation difference of the first translation field is generated;
[0152] The second translation guidance is obtained by combining the translation differences of expression, semantic translation and logical translation.
[0153] Optionally, the target language usage scenario further includes proper noun translation indicators, housing information translation indicators, and apartment structure translation indicators, and the translation guidance also includes third translation guidance. The second translation module 30 is further configured to:
[0154] When the second preset large model identifies that the translation of the proper noun in the text field by the first translation field does not conform to the standard expression of proper nouns in the real estate field, it is determined that the first translation field does not conform to the proper noun translation index, and a proper noun translation difference of the first translation field is generated;
[0155] When the second preset large model identifies that the translation of the property information in the text field by the first translation field does not conform to the standard expression of information in the real estate field, it is determined that the first translation field does not conform to the property information translation index, and a property information translation difference of the first translation field is generated;
[0156] When the second preset large model identifies that the translation of the apartment structure in the text field by the first translation field does not conform to the standard expression of apartment structure in the real estate field, it is determined that the first translation field does not conform to the apartment structure translation index, and a translation difference of the apartment structure for the first translation field is generated;
[0157] The third translation guidance is obtained by combining the translation differences of proper nouns, housing information and apartment structure.
[0158] Optionally, the second translation module 30 is further configured to:
[0159] Integrate the second translation field into the translation text;
[0160] Adjust the corresponding part of the original property information text in the translated text according to the preset field dictionary to obtain the target property information text.
[0161] Optionally, the field adjustment module 40 is further configured to:
[0162] Extract system fields from the original property information text;
[0163] Based on the system field, a match is performed in the preset field dictionary to obtain the standard translation field corresponding to the system field, and the corresponding part of the system field in the translation text is replaced with the standard translation field.
[0164] The text translation device provided in this application utilizes the text translation method described in the aforementioned embodiments to address the technical issue of poor translation accuracy. Compared to the prior art, the text translation device provided in this application achieves the same beneficial effects as the text translation method described in the aforementioned embodiments. Other technical features of the text translation device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0165] The present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the text translation method in the above-mentioned embodiment 1.
[0166] Reference below Figure 5 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, PADs (Portable Application Description: tablet computers), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0167] like Figure 5As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, microphone, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speakers, or vibrator; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wired to exchange data. Although the figures show electronic devices with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.
[0168] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0169] The electronic device provided in this application, employing the text translation method of the above-described embodiment, can resolve the technical problem of poor translation accuracy. Compared to the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the text translation method provided in the above-described embodiment, and the other technical features of the electronic device are the same as those disclosed in the method of the previous embodiment, and are not further described here.
[0170] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0171] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0172] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the text translation method in the above embodiment.
[0173] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0174] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0175] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the text translation device: identifies the original property information text in the property information picture; segments the original property information text to obtain text fields, and translates the text fields into the target language through a first preset large model to obtain a first translation field; detects whether the first translation field conforms to the usage scenario of the target language through a second preset large model, obtains translation guidance, and modifies the text field based on the translation guidance to obtain a second translation field; and obtains the target property information text based on the second translation field.
[0176] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0177] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0178] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0179] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned text translation method, thereby resolving the technical issue of poor translation accuracy. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the text translation method provided in the aforementioned embodiments and are not further elaborated here.
[0180] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned text translation method when executed by a processor.
[0181] The computer program product provided in this application can solve the technical problem of poor translation accuracy. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the text translation method provided in the above embodiment, and will not be repeated here.
[0182] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A text translation method, characterized in that: The text translation method comprises: Detecting text blocks in the house listing image, analyzing the relative positions and semantic relationships of the text blocks to determine the layout rules of the text blocks; Extracting text content from the text block using optical character recognition technology, and adjusting the logical order of the text content in combination with the typesetting rules to obtain the original property information text; Segmenting the original property information text to obtain text fields, and translating the text fields into a target language using a first preset large model to obtain a first translation field; The first translation field is detected by a second preset large model to determine whether it conforms to the usage scenario of the target language, and a translation guide is obtained. The text field is translated based on the translation guide to obtain a second translation field, and the target property information text is obtained based on the second translation field. The target language usage scenario includes a proper noun translation index, a housing information translation index, and a unit structure translation index. The target language usage scenario also includes an expression translation index, a semantic translation index, and a logical translation index. The translation guidance includes a second translation guidance. The step of detecting whether the first translation field conforms to the target language usage scenario through the second preset large model and obtaining the translation guidance includes: detecting and classifying the politeness level of the first translation field using a politeness level analyzer of a second preset large model, and analyzing the first translation field based on politeness expression standards of the target language; if it is detected that the politeness level of the first translation field does not meet the standard or does not conform to the politeness expression standards of the target language, determining that the first translation field does not conform to the expression translation indicator, and generating an expression translation difference for the first translation field; Comparing the semantic similarity between the first translation field and the idiom or idiom using a second preset large model, and if the semantic similarity does not reach a preset similarity threshold or the translation of the idiom or idiom in the text field by the first translation field does not have a matching translation in a preset semantic imagery correspondence table, determining that the first translation field does not meet the semantic translation indicator, and generating a semantic translation difference for the first translation field; Calculating the deviation of the first translation field from the normal word order using the second preset large model, and if the deviation is higher than a preset deviation threshold, determining that the first translation field has disordered word order; or identifying whether there is contradictory information in the first translation field using the second preset large model, and if there is contradictory information, determining that the first translation field is unclear; if the word order of the first translation field is disordered or the meaning is unclear, determining that the first translation field does not meet the logical translation indicator, and generating a logical translation difference for the first translation field; A second translation guidance opinion is obtained by combining the expression translation difference, the semantic translation difference and the logical translation difference.
2. The text translation method according to claim 1, wherein: The target language usage scenario further includes a katakana translation index, a sentence structure translation index, and a unit translation index. The translation guidance also includes a first translation guidance. The step of detecting whether the first translation field conforms to the target language usage scenario by using a second preset large model and obtaining the translation guidance further includes: If the second preset large model identifies that the translation of the katakana in the text field by the first translation field is not a transliteration or a corresponding translated name in the target language, determining that the first translation field does not meet the katakana translation index, and generating a katakana translation difference for the first translation field; Identifying the sentence structure of the first translation field using the second preset large model, and if a component of the sentence structure is missing and affects the meaning of the first translation field, determining that the first translation field does not meet the sentence structure translation indicator, and generating a sentence structure translation difference for the first translation field; If the second preset large model identifies that the translation of the unit in the text field by the first translation field does not conform to the corresponding unit in the target language, determining that the first translation field does not conform to the unit translation indicator, and generating a unit translation difference for the first translation field; A first translation guidance opinion is obtained by combining the katakana translation difference, the sentence structure translation difference and the unit translation difference.
3. The text translation method according to claim 1, wherein: The translation guidance further includes a third translation guidance, and the step of detecting whether the first translation field conforms to the usage scenario of the target language by using the second preset large model, and obtaining the translation guidance further includes: If the second preset large model identifies that the translation of the proper noun in the text field by the first translation field does not conform to the standard expression of proper nouns in the real estate field, determining that the first translation field does not conform to the proper noun translation index, and generating a proper noun translation difference for the first translation field; If the second preset large model identifies that the translation of the property information in the text field by the first translation field does not conform to the standard expression of information in the real estate field, determining that the first translation field does not conform to the property information translation index, and generating a property information translation difference for the first translation field; When the second preset large model identifies that the translation of the apartment structure in the text field by the first translation field does not conform to the standard expression of apartment structure in the real estate field, determining that the first translation field does not conform to the apartment structure translation index, and generating an apartment structure translation difference for the first translation field; A third translation guidance opinion is obtained by combining the translation differences of the proper nouns, the translation differences of the housing information, and the translation differences of the apartment structure.
4. The text translation method according to claim 1, wherein: The step of obtaining the target property information text according to the second translation field includes: Integrating the second translation field into a translation text; The corresponding portion of the original property information text in the translated text is adjusted according to a preset field dictionary.
5. The text translation method according to claim 4, wherein: The step of adjusting the corresponding portion of the original property information text in the translated text according to the preset field dictionary includes: Extracting system fields from the original property information text; A match is performed in a preset field dictionary based on the system field to obtain a standard translation field corresponding to the system field, and a portion corresponding to the system field in the translation text is replaced with the standard translation field.
6. An electronic device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text translation method according to any one of claims 1 to 5.
7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the text translation method according to any one of claims 1 to 5 are implemented.
8. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the text translation method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Term recognition method for multi-language translation
CN116822517A
Batch translation and verification method and device based on large language model, electronic equipment, computer readable storage medium and program product
CN118607539A