Sample generation method and related device
By building virtual scenes with a game engine and automatically extracting text information, the problem of time-consuming and labor-intensive generation of OCR model training samples is solved, and efficient generation of training samples with stable annotation quality is achieved, thereby improving the training efficiency and effectiveness of the OCR model.
Patent Information
- Application Number
- CN202410309713.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-19
Smart Images

Figure CN120673428A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a sample generation method and related devices. Background Art
[0002] Optical Character Recognition (OCR) is a technology that analyzes and recognizes text in image files to obtain information about the text and its layout. In practice, OCR technology is widely used in various scenarios, including transportation, education, office work, finance, and gaming.
[0003] Currently, OCR technology is typically implemented based on an OCR model. This involves feeding an image into a pre-trained OCR model, which then analyzes and identifies the image to obtain text and text layout information. In related technologies, training samples for the OCR model are typically obtained by manually annotating text and text layout information within an image. This sample generation method is time-consuming and labor-intensive, resulting in a limited number of annotated training samples and inconsistent annotation quality. This in turn leads to low training efficiency and poor training results for the OCR model. Summary of the Invention
[0004] The embodiments of the present application provide a sample generation method and related devices, which can quickly obtain a large number of training samples for the OCR model with stable annotation quality, thereby helping to improve the training efficiency and training effect of the OCR model.
[0005] A first aspect of the present application provides a sample generation method, the method comprising:
[0006] Constructing a target virtual scene using a game engine; the target virtual scene includes a model element with text displayed on the surface;
[0007] Rendering the target virtual scene by the game engine to obtain a target scene image;
[0008] During the rendering of the target virtual scene, running a text information extraction script to extract text on the model elements in the target virtual scene and determine position information corresponding to the text in the target scene image;
[0009] According to the extracted text and the corresponding position information, a label corresponding to the target scene image is determined; and the target scene image and the corresponding label are determined to be training samples of an optical character recognition (OCR) model.
[0010] A second aspect of the present application provides a sample generation device, the device comprising:
[0011] A scene construction module is used to construct a target virtual scene through a game engine; the target virtual scene includes a model element with text displayed on the surface;
[0012] An image rendering module, configured to render the target virtual scene through the game engine to obtain a target scene image;
[0013] An information extraction module is used to run a text information extraction script during the rendering of the target virtual scene to extract the text on the model elements in the target virtual scene and determine the position information corresponding to the text in the target scene image;
[0014] The sample determination module is used to determine the label corresponding to the target scene image based on the extracted text and the corresponding position information; and determine that the target scene image and the corresponding label are training samples of the optical character recognition (OCR) model.
[0015] A third aspect of the present application provides a computer device, the device comprising a processor and a memory:
[0016] The memory is used to store computer programs;
[0017] The processor is configured to execute the steps of the sample generation method described in the first aspect according to the computer program.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the steps of the sample generation method described in the first aspect.
[0019] In a fifth aspect, the present application provides a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the sample generation method described in the first aspect.
[0020] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0021] An embodiment of the present application provides a sample generation method, which innovatively utilizes a game engine to construct a training sample for an OCR model. Specifically, in this method, a target virtual scene is first constructed by a game engine, and the target virtual scene includes a model element with text displayed on the surface. Then, the target virtual scene is rendered by the game engine to obtain a target scene image. In order to determine the text and its position information in the target scene image, a pre-written text information extraction script can be run during the rendering of the target virtual scene. The text on the model element in the target virtual scene is extracted by the text information extraction script, and the position information corresponding to the text in the target scene image is determined. Thus, the extracted text and its corresponding position information can be used as labels for the above-mentioned target scene image, and then the target scene image and its corresponding label can be used as training samples for the OCR model. In the embodiment of the present application, with the help of the excellent scene building function and image rendering function of the game engine, a large number of target scene images with rich content style and high quality can be efficiently generated; at the same time, compared with the related art of manually annotating the text and its position information in the image, the present application can use the text information extraction script supported by the game engine to automatically extract text and determine the text position information. Without any manual operation, the text and its position information in the target scene image can be determined efficiently and accurately, thereby obtaining the label corresponding to the target scene image; thereby, a large number of training samples of the OCR model with stable annotation quality can be quickly obtained, and the OCR model can be trained based on these training samples, which is conducive to improving the training efficiency and training effect of the OCR model. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A schematic diagram of an application scenario of the sample generation method provided in an embodiment of the present application;
[0023] Figure 2 A flow chart of a sample generation method provided in an embodiment of the present application;
[0024] Figure 3 A schematic diagram of constructing a target virtual scene provided in an embodiment of the present application;
[0025] Figure 4 A schematic diagram of a training sample for an OCR model provided in an embodiment of the present application;
[0026] Figure 5 A schematic diagram of a diffusion scene image provided in an embodiment of the present application;
[0027] Figure 6 A flowchart of an OCR model training process provided in an embodiment of the present application;
[0028] Figure 7A flowchart of another OCR model training process provided in an embodiment of the present application;
[0029] Figure 8 A schematic diagram of the structure of a sample generation device provided in an embodiment of the present application;
[0030] Figure 9 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;
[0031] Figure 10 A schematic diagram of the structure of the server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0033] The terms "first," "second," "third," "fourth," etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatus.
[0034] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0035] Computer vision (CV) is the science of enabling machines to "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying and measuring objects, and then performing further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0036] Natural language processing (NLP) is a key area of research in computer science and artificial intelligence. It studies theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, the language we use daily, and is closely related to linguistics. It also involves computer science and mathematics, and is a key technology for model training in artificial intelligence. Pre-trained models are derived from large language models in the NLP field. After fine-tuning, large language models can be widely applied to downstream tasks. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question-answering, knowledge graphs, and other technologies.
[0037] The solutions provided in the embodiments of this application involve artificial intelligence computer vision technology and natural language processing technology, which are specifically described through the following embodiments:
[0038] The sample generation method provided in the embodiments of the present application can be executed by a computer device, which can be a terminal device or a server. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server.
[0039] It should be noted that the information, data, and signals involved in the embodiments of this application are authorized by the relevant objects or fully authorized by all parties, and the collection, use, and processing of relevant data comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0040] To facilitate understanding of the sample generation method provided in the embodiments of the present application, the following executor of the sample generation method is used as an example to executor a terminal device to exemplify the application scenario of the sample generation method.
[0041] See also Figure 1 , Figure 1 Schematic diagram of the application scenario of the sample generation method provided in the embodiment of this application. Figure 1 As shown, the application scenario includes a terminal device 110 and a database 120. The terminal device 110 can access the database 120 via a network, or the database 120 can be an embedded database deployed locally on the terminal device 110. The terminal device 110 runs a game engine for executing the sample generation method provided in the embodiment of the present application to construct training samples for the OCR model; the database 120 is used to store the model elements required for constructing the virtual scene, as well as pre-generated text that can be deployed in the virtual scene.
[0042] In actual applications, the terminal device 110 can initiate a data acquisition request to the database 120 to obtain the model elements required for constructing the virtual scene and the text that can be deployed in the virtual scene from the database 120. Furthermore, the terminal device 110 can use the game engine running therein to construct the target virtual scene using the model elements and text obtained from the database 120, and ensure that the constructed target virtual scene includes model elements with text displayed on their surfaces.
[0043] Then, the terminal device 110 can render the target virtual scene through the game engine to obtain the target scene image. In order to determine the text and its position information in the target scene image, a pre-written text information extraction script can be run during the rendering of the target virtual scene to extract the text on the model elements in the target virtual scene and determine the position information corresponding to the text in the target scene image. The extracted text and its corresponding position information can then be used as the label of the above-mentioned target scene image. Compared with the related art of manually annotating the text and its position information in the image, the present application can use the text information extraction script supported by the game engine to automatically extract text and determine the text position information. Without any manual operation, the text in the target scene image and its corresponding position information can be determined efficiently and accurately, thereby obtaining the label corresponding to the target scene image.
[0044] Finally, the terminal device 110 can use the target scene image and its corresponding label as training samples for the OCR model, thereby quickly obtaining a large number of training samples for the OCR model with stable annotation quality; training the OCR model based on these training samples is conducive to improving the training efficiency and training effect of the OCR model.
[0045] It should be understood that Figure 1 The application scenarios shown are only examples. In actual applications, the sample generation method provided in the embodiments of the present application can also be applied to other scenarios. No limitation is imposed on the application scenarios of the sample generation method provided in the embodiments of the present application.
[0046] The sample generation method provided in this application is described in detail below through a method embodiment.
[0047] See also Figure 2 , Figure 2 This is a flow chart of the sample generation method provided in the embodiment of the present application. For the sake of convenience, the following description will be given by taking the execution subject of the sample generation method as an example of a terminal device. Figure 2 As shown, the sample generation method includes the following steps:
[0048] S201: Constructing a target virtual scene through a game engine, where the target virtual scene includes model elements with text displayed on their surfaces.
[0049] A game engine is a core software application used to build video games. It can be a set of precompiled libraries or a fully functional software written specifically for a specific game or game genre, used to develop and create electronic games and other interactive virtual experiences. As an example, a game engine can be a cross-platform game engine such as Unity, Unreal Engine, or Cry Engine. Of course, other types of game engines can also be used, and this application is not limited to a specific game engine.
[0050] Model elements refer to various elements that can be used when constructing a virtual scene, such as characters, buildings, plants, animals, vehicles, and props. Furthermore, depending on the dimensionality of the virtual scene being constructed, model elements of corresponding dimensions can be selected and deployed within the virtual scene. For example, when constructing a 3D virtual scene, 3D model elements can be selected and deployed within the scene. It should be understood that a 3D model element is a mathematical description of the surface of an object in a 3D space, typically represented using 3D computer graphics.
[0051] The target virtual scene refers to a virtual scene constructed based on a game engine, which can be a three-dimensional virtual scene, a two-dimensional virtual scene, etc. In an embodiment of the present application, the above-mentioned target virtual scene can be used as the basic data of the OCR model training sample. Accordingly, the target virtual scene needs to include model elements with text displayed on the surface. For example, the target virtual scene may include billboard elements with advertising slogans displayed on the surface, business buildings with business names displayed on the surface, etc. It should be understood that in actual applications, the target virtual scene may include one or more model elements with text displayed on the surface. The embodiment of the present application does not impose any limitation on the number of model elements for carrying text included in the target virtual scene.
[0052] In an embodiment of the present application, a target virtual scene can be constructed using the virtual scene construction function provided by the game engine. In the process of constructing the target virtual scene, model elements to be deployed in the target virtual scene can be selected from a number of pre-imported model elements. For example, building model elements, road model elements, vehicle model elements, sign model elements, etc. can be selected and deployed in the target virtual scene. In addition, for the model elements deployed in the target virtual scene, relevant text can be set on their surfaces to ensure that text information exists in the constructed target virtual scene, thereby ensuring that the target scene image subsequently obtained based on the target virtual scene includes text information, which can be used as a training sample for the OCR model.
[0053] In a possible implementation, the above-mentioned steps of “building a target virtual scene through a game engine” may include:
[0054] An initial virtual scene is created through a game engine, and then three-dimensional model elements can be deployed in the initial virtual scene. Finally, a target model element can be selected from the deployed three-dimensional model elements, and a text object carrying text can be deployed on the surface of the target model element to obtain a target virtual scene.
[0055] The initial virtual scene refers to a blank virtual scene created based on a game engine; alternatively, the initial virtual scene can be a specific type of template virtual scene, for example, a template virtual scene corresponding to a certain environment (such as a template virtual scene for an urban environment, a template virtual scene for a coastal environment, a template virtual scene for a desert environment, etc.), which includes the basic model elements of the corresponding environment. The target model element refers to the model element in the target virtual scene where text is to be deployed.
[0056] After the initial virtual scene is created through the game engine, three-dimensional model elements can be deployed in the initial virtual scene according to actual needs. For example, three-dimensional model elements such as buildings, roads, traffic signs, trees, vehicles, pedestrians, etc. can be deployed in the initial virtual scene according to needs.
[0057] After deploying the three-dimensional model elements in the initial virtual scene, a model element that can carry text can be selected from the deployed three-dimensional model elements as the target model element; then, a text object is created on the surface of the selected target model element, and the corresponding text content is added to the text object. For example, in the initial virtual scene of an urban environment, three-dimensional model elements such as buildings, roads, traffic signs, trees, vehicles, and pedestrians are deployed. Then, a target three-dimensional model element suitable for deploying text can be selected from these three-dimensional model elements, such as buildings, traffic signs, and vehicles as target three-dimensional model elements. Then, a text object carrying a business name can be deployed on the building, a text object carrying traffic information can be deployed on the traffic sign, a text object carrying a license plate number can be deployed on the vehicle, and so on, thereby obtaining the target virtual scene. Among them, the text deployed on the surface of the target model element can be randomly generated by a script to obtain a variety of different text samples, or it can be obtained from a preset candidate text library. In this regard, this application does not specifically limit the method of obtaining text.
[0058] Taking the game engine Unity as an example, you can first install and open the Unity engine, and then create a new project to store the relevant data of the target virtual scene to be built. After the creation is completed, you can create a new scene in the project view as the initial virtual scene. In addition, you can also import three-dimensional model elements that can be deployed in the initial virtual scene, for example, obtain buildings, billboards, and street signs from the 3D model element database, or you can also obtain the required model elements from 3D modeling software such as the Uinty Asset Store, Blender, and Maya.
[0059] Then, in the scene editor, you can drag the imported model elements into the initial virtual scene. You can adjust the position of the model elements, scale them, and rotate them to create a more realistic virtual scene. For example, if the imported model elements are buildings, billboards, and street signs, you can arrange the buildings into a street in the virtual scene and then add billboards and street signs on both sides of the street.
[0060] Finally, you can select the target model element to which you want to add text, create a new text object on the surface of the target model element, and add the corresponding text content to the text object. For example, you can create a new text object on the surface of a billboard and enter text of the advertising type. In this way, you can obtain the target virtual scene. At the same time, the Unity engine has a rich set of text sample settings built in. By adjusting the font, size, color, and layout of the text, or through the material and mapping functions provided by the Unity engine, you can add various visual effects to the text to obtain more and richer target virtual scenes.
[0061] Therefore, an initial virtual scene is created through a game engine, three-dimensional model elements are deployed in the created initial virtual scene, and target model elements are determined among the three-dimensional model elements deployed in the initial virtual scene. At the same time, text objects carrying text are deployed on the surface of the target model elements, which provides reliable technical support for constructing the target virtual scene and ensures that a richer and higher-quality target virtual scene is obtained.
[0062] In a possible implementation, the above-mentioned step of “deploying three-dimensional model elements in the initial virtual scene” may include:
[0063] Run the scene layout script to deploy 3D model elements in the initial virtual scene according to preset model element deployment rules.
[0064] The implementation steps of “selecting a target model element from the deployed three-dimensional model elements, and deploying a text object carrying text on the surface of the target model element to obtain a target virtual scene” may include:
[0065] Run the scene layout script to select the target model element from the deployed three-dimensional model elements according to the preset model selection rules, and select the target text that matches the target model element from the pre-built candidate text library according to the preset text deployment rules, and deploy the text object carrying the target text on the surface of the target model element to obtain the target virtual scene.
[0066] A scene layout script refers to a script used to automatically layout a virtual scene. Specifically, it can layout model elements in the virtual scene, select target model elements to deploy text from the arranged model elements, and deploy corresponding text on the selected target model elements. Before building the target virtual scene, the scene layout script can be written according to actual needs. Specifically, the procedural modeling function provided by the Unity engine can be used to write the scene layout script. The scene layout script can be written in C or Python. This application does not specifically limit the programming language of the scene layout script. Then, the scene layout script is run during the process of arranging the initial virtual scene.
[0067] The above scene layout script includes preset model element deployment rules, which refer to the rules that record the model element deployment plan pre-set in the scene layout script and are the rules to be referred to when deploying model elements in the virtual scene. Figure 3 , Figure 3 The schematic diagram of constructing the target virtual scene provided in the embodiment of the present application can deploy three-dimensional model elements in the initial virtual scene according to the model element deployment rules preset in the scene layout script. For example, the preset model element deployment rules can be to automatically generate and deploy buildings of random heights on both sides of the street, and then randomly add billboards and street signs on the surface of the buildings. Thus, by running the scene layout script, three-dimensional model elements such as buildings, billboards, and street signs can be deployed in the initial virtual scene according to the model element deployment rules preset in the scene layout script.
[0068] The above-mentioned scene layout script includes preset model selection rules, which are rules used as a reference when selecting the target model element that needs to carry text from the deployed model elements. After the three-dimensional model elements are deployed in the initial virtual scene using the model element deployment rules in the scene layout script, the scene layout script can be continued to run, and the target model element can be selected from the three-dimensional model elements deployed in the initial virtual scene according to the preset model selection rules in the scene layout script. For example, the preset model selection rules can select based on the type of model element, such as selecting a billboard-type model element as the target model element.
[0069] The scene layout script also includes preset text deployment rules. These rules refer to the rules that must be referred to when deploying text on the target model element, specifically indicating the rules for selecting text from a pre-built candidate text library. The pre-built candidate text library contains various types of text, such as advertising-type text, license plate number-type text, and road sign-type text. The preset text deployment rules can be to select the corresponding type of text as the target text based on the target model element. For example, when the target model element is a billboard, advertising-type text can be selected as the target text; when the target model element is a license plate, license plate number-type text can be selected as the target text, and so on. Thus, after determining the target model element in the initial virtual scene through the model selection rules in the scene layout script, the target text that is compatible with the target model element can be selected from the pre-built candidate text library according to the preset text deployment rules, so that a text object carrying the target text can be deployed on the surface of the target model element. For example, a text box containing license plate number-type text can be deployed on the surface of the license plate.
[0070] Therefore, by running the scene layout script, it is possible to automatically layout the initial virtual scene, deploy three-dimensional model elements in the initial virtual scene, select the target model elements that need to carry text, and select the specific text deployed in the target model elements to form the target virtual scene. Through the above method, a large number of rich target virtual scenes can be obtained more quickly, thereby improving the degree of automation of the virtual scene construction process.
[0071] S202: Rendering the target virtual scene through a game engine to obtain a target scene image.
[0072] In computer graphics, rendering is a process of generating images, mainly the process of calculating and generating images from models or other provided information.
[0073] The target scene image is obtained by rendering the target virtual scene using the rendering function of the game engine. It can be used as the basic data for OCR model training samples. For example, the target virtual scene can be rendered using the rendering function of the Unity engine to obtain the target scene image.
[0074] In an embodiment of the present application, the constructed target virtual scene can be rendered using the rendering function provided by the game engine. For example, the target virtual scene can be rendered from a specific perspective and based on specific lighting conditions to obtain a corresponding target scene image. Because the target virtual scene includes model elements with text displayed on the surface, the target scene image obtained by rendering the target virtual scene will also reflect the model elements with text displayed on the surface, thereby ensuring that the rendered target scene image includes text and can be used as a training sample for the OCR model.
[0075] In a possible implementation, the above-mentioned steps of “rendering the target virtual scene through a game engine to obtain a target scene image” may include:
[0076] Determine rendering scene parameters, which are used to determine the lighting conditions corresponding to the rendered scene image, the viewing angle corresponding to the scene image in the target virtual scene, and at least one of the materials and textures of the model elements in the scene image; through the game engine, the target virtual scene is rendered based on the rendering scene parameters to obtain the target scene image.
[0077] Rendering scene parameters refer to the parameters based on which the rendering function in the game engine performs rendering processing. The rendering scene parameters may include, for example: lighting parameters for determining the lighting conditions corresponding to the rendered scene image, camera parameters for determining the viewing angle corresponding to the rendered scene image in the target virtual scene, material parameters for determining the material of the model elements in the rendered scene image, and texture parameters for determining the texture of the model elements in the rendered scene image, etc. By adjusting the rendering scene parameters, different rendering effects can be obtained, thereby obtaining realistic images under different environmental conditions.
[0078] As an example, taking the game engine Unity as an example, by adjusting the lighting parameters, camera parameters, etc. of the Unity engine, the target virtual scene can be rendered from different perspectives to generate realistic images with different lighting and perspectives.
[0079] The lighting reference in the Unity engine includes various light source type parameters (directional light type, point light type and spotlight type) and global illumination system parameters (light map and precomputed real-time global illumination). By adjusting the above lighting parameters, you can simulate the lighting effects of the real world.
[0080] Specifically, one or more light sources can be added to the target virtual scene. By adjusting parameters such as the light source type parameters, color parameters, intensity parameters, and light source direction parameters of each light source, various lighting conditions can be simulated in the rendered scene image, such as daylight, indoor lighting, and shadows.
[0081] Then, the overall lighting effect of the rendered scene image can be changed by adjusting the global illumination system parameters, such as the skybox parameters (i.e., the parameters used to create the natural environment in the image), the ambient light color parameters and the ambient light intensity parameters, and the global illumination mode parameters, to create lighting effects under different time and weather conditions, so that the target scene image obtained by the rendering function is closer to the real environment, thereby enhancing the realism and atmosphere of the target scene image.
[0082] In addition, you can adjust camera parameters to obtain scene images from different perspectives. In the Unity engine, the camera is the window for observing and rendering the scene. Camera parameters include camera position parameters, angle parameters, and lens parameters (focal length and field of view). By adjusting these camera parameters, you can simulate real-world visual effects.
[0083] Specifically, one or more cameras can be added to the target virtual scene, each capable of observing the scene from a different position and angle. By adjusting the camera's position parameters and the angle parameters by rotating or inputting numerical values, the viewpoint of the scene can be changed. This means the camera's position and viewing angle, i.e., the camera's field of view and projection method, can be changed. For example, a camera can be placed at eye level to simulate a human perspective, or placed in the air to simulate a bird's-eye view.
[0084] Then, you can change the field of view and perspective of the scene by adjusting the camera lens parameters, such as focal length and field of view. For example, you can set a wide-angle lens to obtain a wide field of view, or set a telephoto lens to simulate a farsighted effect.
[0085] In addition, the material and texture of the model elements in the rendered scene image can be changed by adjusting the material parameters and texture parameters. For example, the material of the model elements in the rendered scene image can be made to look like metal, wood, plastic, etc., and the texture of the model elements in the rendered scene image can have different surfaces such as smooth, rough, and glossy, so as to obtain a scene image with more stable quality and richer quality.
[0086] Once the rendering scene parameters are determined, the target virtual scene can be rendered based on the rendering scene parameters using the game engine's rendering function to obtain the target scene image. As an example, you can select the camera you want to render in the Unity engine, that is, select the target virtual scene with a specific perspective you want to render. Then, you can select the view window "Game" in the menu bar to display the view of the currently selected camera. Then, click the "Play" button in the "Game" window to start rendering the target virtual scene.
[0087] During the rendering process, the Unity engine calculates the color of each pixel in the target virtual scene based on the rendering scene parameters, namely at least one of the following parameters: lighting parameters, camera parameters, material parameters, and texture parameters. Finally, the Unity engine saves the calculation results as an image text, for example, a two-dimensional image text, namely the target scene image.
[0088] Therefore, by adjusting the rendering scene parameters, scene images under various lighting and viewing angles can be obtained. By rendering the target virtual scene based on the rendering scene parameters through the rendering function of the game engine, target scene images that better meet actual needs and have richer image content can be obtained. Training the OCR model based on this target scene image is conducive to improving the robustness and generalization ability of the OCR model, and is conducive to improving the adaptability of the OCR model to images under various environmental conditions.
[0089] S203: During the process of rendering the target virtual scene, a text information extraction script is run to extract text on model elements in the target virtual scene and determine position information corresponding to the text in the target scene image.
[0090] A text information extraction script is a script used to extract text and its corresponding location information from a target scene image. The text information extraction script can be pre-written and created, and then run during the rendering of the target virtual scene. The text information extraction script can be written in C or Python, and this application does not specifically limit the programming language of the text information extraction script.
[0091] During the rendering of the target virtual scene, running the text information extraction script can automatically extract text from the model elements of the target virtual scene and obtain the position information corresponding to the text in the target scene image.
[0092] In one possible implementation, the steps of “running a text information extraction script to extract text on model elements in the target virtual scene and determining position information corresponding to the text in the target scene image” may include:
[0093] Run the text information extraction script to obtain the text on the model elements in the target virtual scene from the text component of the game engine; and determine the positioning point corresponding to the display area of the text through the positioning component in the text component, convert the positioning point into the corresponding screen coordinates, and obtain the corresponding position information of the text.
[0094] As an example, taking the Unity engine as an example, in the Unity engine, a scripting language (such as C#) can be used to create a text information extraction script. The text information extraction script is used to record the text content in the rendered virtual scene and its position information in the scene image, and output both.
[0095] Specifically, when running the text information extraction script, the text content can be obtained through the text component Text in the Unity engine. In addition, the positioning component RectTransform in the Text component in the Unity engine can be used to calculate the positioning point corresponding to the display area of the text, that is, the positioning point of the text content displayed in the scene image can be calculated. Afterwards, the coordinate conversion function RectTransformUtility.ScreenPointToLocalPointInRectangle in the Unity engine can be used to convert the positioning point corresponding to the text content into screen coordinates, that is, the four positioning points of the text content located in the scene image can be converted into corresponding screen coordinates through the coordinate conversion function to obtain the position information corresponding to the text. Furthermore, the extracted text and the corresponding position information can be output to a file as a label for the target scene image. In addition, the format of the output label can be any format acceptable to the machine learning model, for example, a plain text file format (Comma-Separated Values, CSV), a lightweight data exchange format (JavaScript Object Notation, JSON) or a plain text format (Extensible Markup Language, XML). In this regard, this application does not specifically limit the format of the output label.
[0096] After obtaining the text information extraction script, you can add a text information extraction script to each model element containing text in the target virtual scene, so that the content of the text on the model element and its position information can be automatically recorded during the rendering process. Afterwards, after starting to render the scene image, the text information extraction script will be run for each text in the target virtual scene, so that the text content on the model element in the target virtual scene and the position information of the text on the target scene image can be automatically recorded. Finally, the text in the target scene image and its corresponding position information can be output to the label file to obtain the label of the target scene image. In addition, after the rendering is completed, the label text can be post-processed according to actual needs, for example, merging multiple label files, or converting the label format to meet the training requirements of the OCR model.
[0097] Therefore, through the above method, in the process of rendering the target virtual scene, by running the text information extraction script, the text in the target scene image and its corresponding position information can be automatically obtained, and the automatic creation and recording of the target scene image label is realized, and the labeling work of the OCR model training samples is efficiently completed, which greatly saves the labeling time and the workload of manual labeling, and at the same time ensures the accuracy and format consistency of the labeled labels.
[0098] S204: Determine a label corresponding to the target scene image based on the extracted text and its corresponding position information; and determine that the target scene image and its corresponding label are training samples for an optical character recognition (OCR) model.
[0099] Finally, the text extracted from the target scene image and its corresponding position information can be determined as the label corresponding to the target scene image. After that, the target scene image and its corresponding label can be used as training samples for the OCR model.
[0100] It should be noted that the label corresponding to the target scene image can be a text-editable file with the same size as the target scene image. The text in the target scene image is displayed in the file, and the display position of the text is consistent with its display position in the target scene image. Therefore, the target scene image containing the corresponding label can be used as a training sample for the OCR model. Figure 4 , Figure 4 Schematic diagram of the training sample of the OCR model provided in the embodiment of the present application. In this regard, the present application does not specifically limit the representation form of the label of the target scene image.
[0101] In the sample generation method provided in the embodiment of the present application, the method innovatively utilizes the game engine to construct the training sample of the OCR model. Specifically, in this method, the target virtual scene is first constructed by the game engine, and the target virtual scene includes a model element with text displayed on the surface, and then the target virtual scene is rendered by the game engine to obtain a target scene image. In order to determine the text and its position information in the target scene image, a pre-written text information extraction script can be run during the rendering of the target virtual scene, and the text on the model element in the target virtual scene can be extracted by the text information extraction script, and the position information corresponding to the text in the target scene image can be determined. Thus, the extracted text and its corresponding position information can be used as the label of the above-mentioned target scene image, and then the target scene image and its corresponding label can be used as the training sample of the OCR model. In the embodiment of the present application, with the help of the excellent scene building function and image rendering function of the game engine, a large number of target scene images with rich content style and high quality can be efficiently generated; at the same time, compared with the related art of manually annotating the text and its position information in the image, the present application can use the text information extraction script supported by the game engine to automatically extract text and determine the text position information. Without any manual operation, the text and its position information in the target scene image can be determined efficiently and accurately, thereby obtaining the label corresponding to the target scene image; thereby, a large number of training samples of the OCR model with stable annotation quality can be quickly obtained, and the OCR model can be trained based on these training samples, which is conducive to improving the training efficiency and training effect of the OCR model.
[0102] In a possible implementation, the above-mentioned “text displayed on the surface of the model element” can be generated in the following manner:
[0103] Determine model reference information, where the model reference information includes at least one of prompt text and model reference parameters, where the prompt text is used to indicate content features corresponding to the text to be generated, and the model reference parameters are used to control basic features corresponding to the output results of the large language model; call the large language model, and generate text displayed on the surface of the model element based on the model reference information.
[0104] The model reference information refers to the information that a large language model (LLM) needs to refer to when the LLM is called to generate text. The model reference information may include, for example, prompt text and model reference parameters.
[0105] The prompt text is used to indicate the content features corresponding to the text to be generated. The content features may specifically include the type of text, the semantic information expressed by the text, etc. For example, the prompt text may be "Generate a random license plate number in the domestic license plate format," "Generate an advertising slogan for a specific industry or product," and so on. In the embodiment of the present application, it is necessary to generate a large amount of diverse text, such as license plate numbers, advertising slogans, text in various languages, and other possible text types, such as book text, street signs, handwritten notes, etc. Accordingly, various different types of text can be obtained by setting corresponding prompt text.
[0106] The model reference parameter is used to control the basic features corresponding to the output results of the large language model. The basic features refer to the basic properties of the text generated by the large language model. For example, the model reference parameter can be temperature, maximum output length or minimum output length. Temperature is a parameter that can adjust the randomness of the model output. A higher temperature value (such as 1.0 or above) will cause the model to generate text with higher randomness (the larger the temperature value, the higher the randomness of the output); a lower temperature value (such as 0.5 or below) will cause the model to generate more conservative text (the smaller the temperature value, the closer the output is to the most likely result considered by the model); thus, the diversity of the generated text can be controlled by adjusting the temperature value of the model. The maximum output length and the minimum output length can control the length of the text generated by the model; in actual applications, the maximum output length and the minimum output length of the model can be set according to the required text length range to ensure that the text generated by the model meets the expected length requirements. For example, if you need to generate a license plate number with a length between 6 and 10, you can set the maximum output length to 10 and the minimum output length to 6 to limit the length of the generated text. Of course, in the embodiment of the present application, the above-mentioned model reference parameters can also include other parameters, and the embodiment of the present application does not make any limitation on this.
[0107] A large language model is a general-purpose text analysis solution that uses a large number of parameters and can process large amounts of text data. Large language models are pre-trained on large amounts of text data and can then be fine-tuned for various tasks. In this embodiment, they are used to generate text deployed in a target virtual scene.
[0108] As an example, before generating text deployed in a target virtual scene based on a large language model, access rights to the application programming interface (API) of the large language model can be obtained to obtain an API key, thereby facilitating subsequent sending of API requests to the large language model based on the API key. The large language model can be a third-generation generative pre-trained model (GPT-3) or other pre-trained models. This application does not specifically limit the large language model.
[0109] When calling a large language model to generate text displayed on the surface of a model element based on model reference information, the text requirements for the text deployed in the target virtual scene can be clarified first, such as the type, style, length, and theme of the text to be generated. Text requirements can cover license plate numbers, advertising slogans, text in various languages, and other possible text types, and the model reference information used when calling the large language model can be set accordingly. For example, one or more prompt texts can be designed for each type of text to guide the model to generate text that meets the text requirements. At the same time, model reference parameters such as the model temperature or maximum output length can be set to control the randomness of the generated text and basic features corresponding to the output results, such as length.
[0110] Specifically, when determining text requirements, it is necessary to determine the text type of the text. For example, when a license plate number is required, it is necessary to consider the license plate formats of different countries and regions; when an advertising slogan is required, it is necessary to consider different product types, such as food, games, cars, electronic products, etc. For each text type, one or more prompt texts can be created. For example, for the license plate type, the prompt text is "include a specific letter or number combination on the license plate", and for the advertising slogan type, the prompt text is "display a specific brand name or logo on the billboard". At the same time, the model reference parameters can also be adjusted according to different text types. For example, for the license plate type, the maximum output length is no more than 7 digits, and for the advertising slogan type, the temperature value is within the first reference value range.
[0111] Finally, the large language model is called based on the API key obtained previously. Specifically, the API key can be entered first. After the API key is verified, an API request carrying the above-mentioned prompt text and model reference parameters is sent to the large language model, so that the large language model generates text based on the prompt text, and controls the basic features corresponding to the model output results based on the model reference parameters to obtain output results that meet the text requirements, that is, text that meets the text requirements and is deployed in the target virtual scene.
[0112] It should be noted that the large language model can be used to generate text when constructing the target virtual scene, or it can be used to pre-generate text before constructing the target virtual scene and store the generated text in the database. Furthermore, when deploying text in the target virtual scene according to preset text deployment rules, the large language model can be used to generate text in advance. The text generated by the large language model can then be classified by text type and stored in a candidate text library. This facilitates subsequent matching and searching based on text type and corresponding model element type, thereby improving the deployment efficiency of model elements.
[0113] In one possible implementation, after the text deployed in the target virtual scene is generated by the large language model, the quality and diversity of the generated text can be evaluated, for example, whether the generated text meets the preset text requirements. If so, the text is considered to have a higher quality; if not, the text is considered to have a lower quality. The preset text requirements refer to the evaluation basis for the text generated by the large language model, which is used to evaluate the quality of the text generated based on the large language model. For example, whether the degree of match between the text generated by the large language model and the model reference information is greater than a first preset threshold, such as calculating the degree of match between the generated text and the prompt text. If the degree of match is greater than the first preset threshold, the quality of the generated text is considered to be better. Otherwise, the quality of the generated text is considered to be poor.
[0114] If the text generated by calling the large language model does not meet the preset text requirements, an information iterative adjustment operation can be performed until text that meets the preset text requirements is obtained. The information iterative adjustment operation includes adjusting the model reference information used in the current call of the large language model, for example, adjusting at least one of the prompt text and the model reference parameters to obtain adjusted model reference information, and then calling the large language model again to regenerate text based on the adjusted model reference information.
[0115] Therefore, when the text generated by calling the large language model does not meet the preset text requirements, the model reference information used when calling the large language model this time can be adjusted to obtain the adjusted model reference information, and then the large language model can be called again to generate text based on the adjusted model reference information until the text that meets the preset text requirements is obtained. By continuously iteratively adjusting the model reference information, the quality of the text can be continuously improved, so that a large amount of high-quality text can be obtained.
[0116] As an example, refer to the following code, which details the process of calling a large language model to generate text.
[0117]
[0118]
[0119] In the above code, the API key is first set to enable subsequent security control and authentication of access to the API. By using the API key, the API provider can ensure that only authorized applications or users can access its API to enhance the security of the text generation process. Then, a function generate_text for generating text is defined. This function uses the creation of the artificial intelligence method openai.Completion.create and uses the API key and the set request parameters to send an API text generation request to OpenAI's server and return the generated text. After obtaining the generated text, the model reference parameters are set in the main program, such as the maximum output length (max_length) and temperature (temperature), and the text type is specified as a license plate number. Finally, the above steps can be repeated to generate a large amount and diversity of text in a loop, and prompt text can be used to generate text with a specific format.
[0120] After each round of text generation, the quality of the generated text can be checked according to the preset text requirements to ensure the quality of the generated text; if the preset text requirements are not met, the large language model can be iteratively adjusted by adjusting the model reference information until text that meets the preset text requirements is generated.
[0121] Finally, the generated text can be classified and stored according to attributes such as text type and language, and stored in the database accordingly, which facilitates subsequent search based on attribute information, thereby improving the efficiency of text deployment.
[0122] Therefore, by setting at least one of the prompt text and model reference parameters, the model reference information can be determined, and then the large language model can be called. Based on the determined model reference information, the text displayed on the surface of the model element can be generated. This method can not only generate a large amount of diverse text, but also provide rich material samples for the subsequent generation of OCR model training samples.
[0123] In a possible implementation, after obtaining the target scene image, the rendered target scene image may be subjected to diffusion processing to obtain richer image samples. The implementation steps of performing diffusion processing on the rendered target scene image may include:
[0124] A stable diffusion algorithm is used to perform diffusion processing on the target scene image to obtain a diffuse scene image; the diffusion processing is used to deform the entire target scene image or a local area in the target scene image; the label corresponding to the diffuse scene image is determined according to the label corresponding to the target scene image; and the diffuse scene image and its corresponding label are determined as training samples for the OCR model.
[0125] The Stable Diffusion algorithm is an image processing algorithm that can deform, distort, and change images. It can change the style of an image while maintaining its basic structure, thereby generating sample images of various styles. For example, it can generate sample images of anime, science fiction, and ancient styles to increase sample diversity. The Stable Diffusion algorithm can be explicit stable diffusion, implicit stable diffusion, or other deformation algorithms based on physical simulation. The embodiments of this application do not limit the specific stable diffusion algorithm.
[0126] After obtaining the target scene image, the stable diffusion algorithm can be applied to the target scene image for diffusion processing. Specifically, refer to Figure 5 , Figure 5 A schematic diagram of a diffuse scene image provided in an embodiment of the present application can perform deformation processing on the entire target scene image, or can selectively perform deformation processing on a local area or text in the target scene image to obtain a diffuse scene image, so that the text or image in the diffuse scene image is deformed, distorted or flows, thereby increasing the diversity of samples.
[0127] The diffused scene image refers to an image obtained by performing diffusion processing on a scene image. For example, the diffused scene image may be an image of the target scene image after overall deformation, an image of the target scene image after local deformation, or an image of text in the target scene image after deformation.
[0128] In a possible implementation, the steps of “using a stable diffusion algorithm to perform diffusion processing on the target scene image to obtain a diffused scene image” may include:
[0129] Determine diffusion processing parameters; the diffusion processing parameters are used to indicate at least one of a diffusion step size, a number of diffusion iterations, a deformation intensity, a deformation effect, and a deformation area range; adopt a stable diffusion algorithm, based on the diffusion processing parameters, perform diffusion processing on the target scene image to obtain a diffused scene image.
[0130] Diffusion processing parameters refer to the parameters used when performing diffusion processing on the target scene image based on the stable diffusion algorithm. Diffusion processing parameters may include, for example, diffusion step size, number of diffusion iterations, deformation intensity, deformation effect, deformation area range, etc. By adjusting the above diffusion processing parameters, the degree of change and diversity of the samples can be controlled.
[0131] The diffusion step size controls the extent of image diffusion during each diffusion operation. A larger step size results in a more pronounced deformation effect, while a smaller step size produces more subtle changes. An appropriate step size can be selected based on the desired deformation. The number of iterations refers to the number of iterations of the stable diffusion algorithm. A higher number of iterations produces a more pronounced deformation effect, while a lower number of iterations produces a more subtle change. This value can be adjusted based on the desired deformation effect and computational resources to ensure the quality of the diffused scene image. The deformation strength controls the extent of the diffusion operation's effect on the image. A higher deformation strength results in greater deformation and distortion, while a lower deformation strength produces less distortion. An appropriate deformation strength can be selected based on the desired deformation. The deformation region range controls the region of the target scene image to which the diffusion operation is applied. Diffusion can be applied to the entire image or to specific regions or text. This can be set based on the desired deformation effect and training requirements. The deformation effect controls the actual effect achieved after the diffusion operation, such as distortion or flow. An appropriate deformation effect can be selected based on the desired deformation effect.
[0132] In addition, the diffusion processing parameters may also be other parameters. For example, depending on the selected stable diffusion algorithm, there may be other adjustable parameters such as velocity factor, deformation function, gradient constraint, etc. The specific settings of these parameters depend on the requirements and effects of the selected stable diffusion algorithm. In this regard, this application does not limit the specific diffusion processing parameters.
[0133] After determining the diffusion processing parameters, a stable diffusion algorithm can be used to perform diffusion processing on the target scene image based on the diffusion processing parameters. During the diffusion processing process, the diffusion effect can be controlled by adjusting the diffusion processing parameters or using other image processing techniques according to the actual deformation requirements, thereby obtaining a diffuse scene image. For example, the diffusion step size, the diffusion speed, the deformation amplitude, or different diffusion algorithms can be adjusted to obtain a more diverse change effect. In addition, different levels of diversity changes can be generated according to the actual deformation requirements. For example, slight deformations and distortions can be generated, or more complex deformations and flow effects can be generated. By adjusting the diffusion processing parameters, such as the number of diffusion iterations and the range of the deformation area, the level of diversity changes can be controlled to meet the training requirements of different OCR models.
[0134] Finally, the generated diffusion scene images may contain some images whose actual deformation effects are too large or do not meet the deformation requirements. For example, the actual deformation effect is too different from the target scene image, or the deformation requirement is an ancient style image, while the generated diffusion scene image is a two-dimensional style image. Therefore, after obtaining the diffusion scene images, data cleaning and screening can be performed first to remove images with poor quality (images in the above cases) or images that are not suitable for model training (there is no obvious text in the deformed diffusion scene images), so that more high-quality diffusion scene images can be obtained, thereby ensuring the sample quality as OCR model training samples.
[0135] As an example, please refer to the code below, which specifically introduces the process of using the stable diffusion algorithm to diffuse the target scene image to obtain the diffused scene image.
[0136]
[0137]
[0138] In the above code, a blank image of the same size as the target scene image is first created as output. The stable diffusion algorithm is then executed, applied to each pixel in the target scene image, and the diffusion change is calculated and applied to the brightness value of the current pixel. The original target scene image is then loaded as a grayscale image, and diffusion processing parameters are set, such as the number of diffusion iterations and the deformation strength. Finally, the stable diffusion algorithm is run on the target scene image to obtain a diffused scene image, which can then be saved to a database.
[0139] Therefore, through the above method, a stable diffusion algorithm can be used to perform diffusion processing on the target scene image based on the diffusion processing parameters. By adjusting the diffusion processing parameters, a more diverse change effect can be obtained, and a large number of rich and diverse diffusion scene images can be obtained, thereby obtaining a large number of rich and diverse training samples for the OCR model.
[0140] After obtaining the diffuse scene image, the label corresponding to the diffuse scene image can be determined according to the label corresponding to the target scene image. Finally, the diffuse scene image and its corresponding label can be used as training samples for the OCR model.
[0141] As an example, when the deformation effect of the text in the diffuse scene image is small (the text in the image can be clearly identified), the label corresponding to the target scene image, that is, the text extracted from the target scene image and its corresponding position information, can be directly used as the label corresponding to the diffuse scene image; when the deformation effect of the text in the diffuse scene image is large (the text in the image cannot be clearly identified), a text information extraction script can be run during the rendering process of the diffuse scene image to extract the text and its corresponding position information from the diffuse scene image, and the extracted text and its corresponding position information can be used as the label corresponding to the diffuse scene image. Alternatively, other text extraction methods can be used to extract text and its corresponding position information from the diffuse scene image. This application does not limit the specific text extraction method.
[0142] Therefore, by combining the above method with the stable diffusion algorithm, the target scene image is diffused, and a large number of diffuse scene images with diverse changes can be obtained. At the same time, based on the label corresponding to the target scene image, the label corresponding to the diffuse scene image can be determined, and the diffuse scene image and its corresponding label can be used as training samples of the OCR model, so that a large number of training samples of the OCR model with stable annotation quality can be obtained. The above method can not only generate more, more diverse and actual-fit OCR model training samples, but also help improve the training effect of the OCR model, and help improve the robustness and generalization ability of the OCR model.
[0143] In one possible implementation, reference may be made to Figure 6 , Figure 6 The following is a flowchart of an OCR model training process provided in an embodiment of the present application. The OCR model can be trained by the following steps:
[0144] S601: Acquire a training sample set, where the training sample set includes training samples generated based on a virtual scene constructed by a game engine.
[0145] An OCR model is a model that analyzes and recognizes image files containing text to obtain text and layout information (i.e., text location information). In actual applications, before a server trains an OCR model, it needs to obtain a set of training samples for training the OCR model. The training sample set includes training samples generated using the sample generation method provided in the embodiments of this application, i.e., training samples generated based on a virtual scene constructed using a game engine.
[0146] In addition, after obtaining the training sample set, the training samples can also be preprocessed, such as performing image normalization, expansion and other operations on the training samples in the training sample set, to facilitate subsequent model training based on the preprocessed training samples, so as to reduce the impact of unnecessary factors on model training.
[0147] S602: Using the OCR model to be trained, the image in the training sample is recognized to obtain predicted text information corresponding to the training sample, where the predicted text information includes the predicted text included in the image and predicted position information corresponding to the predicted text.
[0148] After obtaining the training sample set, the server can input the training samples in the training sample set into the OCR model to be trained. After the OCR model recognizes the image in the input training sample, it can output the predicted text information corresponding to the training sample. The predicted text information refers to the output result of the OCR model to be trained, which is used to characterize the text predicted by the trained OCR model for the image in the training sample (i.e., predicted text), and the predicted position information corresponding to the text (i.e., predicted position information). As an example, the predicted text and its corresponding predicted position information can be displayed in the predicted text information in the form of a labeling box in the image. In this regard, the embodiment of the present application does not specifically limit the form of expression of the predicted text and its corresponding predicted position information in the image.
[0149] It should be understood that the predicted text is the text determined by the trained OCR model from the training sample, and the predicted position information is the position of the text determined by the trained OCR model in the training sample. At the same time, this application does not specifically limit the model structure of the OCR model.
[0150] S603: Training an OCR model based on the labels of the training samples in the training sample set and the corresponding predicted text information.
[0151] The labels included in each training sample in the training sample set contain the text corresponding to the image in each training sample and its corresponding position information. After obtaining the predicted text information corresponding to the training sample, the difference between the predicted text information and the labels included in each training sample can be calculated through the loss function, that is, the difference between the predicted text in the predicted text information and the text in the label is calculated, and the difference between the predicted position information in the predicted text information and the text position information in the label is calculated. Therefore, the OCR model can be trained based on the calculated loss value, that is, with the goal of reducing the loss value, the recognition performance of the OCR model can be optimized by continuously adjusting the parameters of the OCR model to improve the accuracy and recall rate of the OCR model. This application does not specifically limit the loss function in the OCR model training process.
[0152] In addition, the trained OCR model can be tested using the validation set. By comparing the differences between the model's predicted text information and the actual text and its corresponding position information in the validation set, the accuracy and recall of the OCR model can be evaluated. When the accuracy or recall of the OCR model is lower than the preset threshold, the OCR model can be further adjusted and optimized. For example, the OCR model parameters can be adjusted, the model structure can be changed, or other optimization strategies can be adopted to improve the recognition performance of the OCR model.
[0153] It should be understood that in actual applications, a training end condition can be pre-set. When the training of the OCR model reaches the training end condition, the training of the OCR model can be considered to be terminated. The training end condition can be, for example, that the number of training rounds for the OCR model reaches a preset round threshold. Another example can be that the performance of the OCR model is tested and it is found that the performance of the OCR model reaches a preset performance standard (such as reaching a preset recognition accuracy, etc.). Another example can be that the performance of the OCR model is tested and it is found that the performance of the OCR model no longer improves significantly with the progress of training. The embodiments of the present application do not impose any restrictions on the training end condition.
[0154] Therefore, the training samples generated based on the virtual scene constructed by the game engine constitute a training sample set, and the final OCR model can be obtained based on the above training steps and the training sample set. Since the training sample set includes a large number of rich and diverse training samples, it helps the trained OCR model to more accurately identify text in various images and its corresponding position information, that is, improve the training effect of the OCR model.
[0155] In one possible implementation, in order to improve the authenticity of the training samples of the generated OCR model, a generative adversarial network (GAN) can be introduced into the training process of the OCR model, that is, the above-mentioned "training sample set" can also include training samples generated based on real scenes. The training samples generated based on real scenes include images based on real scenes, such as images obtained by photographing real scenes. The training samples generated based on real scenes also include labels corresponding to the images, including text in the image and its corresponding location information. The labels can be obtained through manual annotation or other methods. The embodiments of the present application do not impose any restrictions on this.
[0156] Correspondingly, refer to Figure 7 , Figure 7 A flowchart of another OCR model training process provided in an embodiment of the present application, wherein the OCR model training steps further include:
[0157] S701: Classify images in training samples using an OCR model to obtain predicted image types corresponding to the training samples. The predicted image types are used to characterize whether the images predicted by the OCR model are generated based on a virtual scene or a real scene.
[0158] GAN is a model structure composed of a generator network and a discriminator network, which compete with each other through adversarial training to improve the quality of generated samples. In an embodiment of the present application, the discriminator in GAN is introduced into the training process of the OCR model to add a branch task for identifying image types in the process of training the OCR model, so that the OCR model can accurately identify both images generated based on real scenes and images generated based on virtual scenes, avoiding the uneven distribution of training samples generated based on real scenes and training samples generated based on virtual scenes (i.e., training samples generated by the game engine) in the training sample set, which causes the OCR model to only accurately identify images in a certain specific scene, while it is difficult to accurately identify images in other scenes.
[0159] After obtaining the training sample set, the training samples in the training sample set can be input into the OCR model, and the OCR model can be used to perform text recognition and image classification based on the images in the training samples. Specifically, the relevant processing of text recognition by the OCR model can be found in S602 above; in addition, a discriminator network can be added to the OCR model, and the images in the training samples can be classified by the discriminator network to determine the predicted image type corresponding to the image in the training sample, that is, to determine whether the image is generated based on a real scene or a virtual scene (that is, based on a target virtual scene generated by a game engine). The predicted image type refers to the prediction result obtained by classifying the images in the training sample by the OCR model, which is used to indicate whether the prediction result is an image generated based on a real scene or an image generated based on a virtual scene constructed by a game engine.
[0160] S702: Construct a first loss function based on the labels of the training samples in the training sample set and the corresponding predicted text information.
[0161] The first loss function is used to measure the difference between the predicted text information and the text information indicated by the label in the training sample, that is, the difference between the predicted text output by the OCR model and the text in the label, and the difference between the predicted position information output by the OCR model and the text position information in the label. Specifically, the first loss function can be constructed based on the difference between the labels included in each training sample in the training sample set and the predicted text information corresponding to each. As an example, the first loss function can be obtained based on the cross entropy between the true label distribution (the labels included in each training sample in the training sample set) and the predicted label distribution (the predicted text information corresponding to each training sample in the training sample set) as the loss. In contrast, the present application does not specifically limit the type of the first loss function.
[0162] S703: Constructing a second loss function based on the predicted image type and the actual image type corresponding to each training sample in the training sample set.
[0163] The second loss function refers to a function that measures the difference between the predicted image type and the actual image type corresponding to the image in the training sample. Specifically, the second loss function can be constructed based on the difference between the actual image type corresponding to each training sample in the training sample set and the predicted image type corresponding to each training sample. Among them, the actual image type refers to the type to which the image in the training sample actually belongs, which is used to indicate whether the image in the training sample is generated based on a real scene or a virtual scene. As an example, the second loss function can be obtained based on the cross entropy between the real label distribution (actual image type) and the predicted label distribution (predicted image type) as the loss. In contrast, the present application does not specifically limit the type of the second loss function.
[0164] S704: Train an OCR model according to the first loss function and the second loss function.
[0165] Finally, the OCR model can be trained using a first loss function constructed based on the difference between the true label and the predicted text information, and a second loss function constructed based on the difference between the actual image type and the predicted image type. As an example, different weight parameters can be assigned to the first loss function and the second loss function, and a comprehensive loss value consisting of a first loss value of the first loss function and a second loss value of the second loss function can be calculated based on the respective weight parameters. The OCR model can be trained to optimize the recognition performance of the OCR model by continuously adjusting the parameters of the OCR model.
[0166] During the training process, the training goal of the discriminator network is to accurately determine the type of input image (i.e., whether it is generated based on a real scene or a virtual scene). In an embodiment of the present application, the training samples generated based on virtual scenes in the training sample set may be far more than the training samples generated based on real scenes, that is, since the method of generating training samples through a game engine is efficient and low-cost, a large number of training samples based on virtual scenes generated by a game engine can be obtained, while the acquisition cost of training samples generated based on real scenes is high, so the training samples generated based on real scenes may be less; when the training sample set includes more training samples generated based on virtual scenes and fewer training samples generated based on real scenes, when the OCR model is trained based on the training sample set, the OCR model will tend to have better recognition performance for images generated based on virtual scenes, while the recognition performance for images generated based on real scenes is relatively poor.
[0167] In order to solve the above problem and avoid the OCR model being able to accurately recognize text information in images generated based on virtual scenes only, the embodiment of the present application proposes the above-mentioned mechanism of introducing the discriminator network in GAN during the training of the OCR model, so that the OCR model can perform the image classification task while performing the text recognition task, and identify whether the input image is generated based on a real scene or a virtual scene. For the image classification task during the model training process, the OCR model tends to confuse the input image, making it difficult to accurately distinguish whether the input image is generated based on a virtual scene or an image generated based on a real scene. For example, for the image classification task, the accuracy of the image classification task tends to be around 50% during the training process; that is, the OCR model has the same processing capabilities for images generated based on virtual scenes and images generated based on real scenes, and does not have a tendency due to the different ways in which the input images are generated.
[0168] Therefore, by introducing the discriminator network in the process of training the OCR model, the OCR model can perform image classification tasks simultaneously during the training process, which can help the OCR model have better text recognition performance for images generated based on real scenes and images generated based on virtual scenes, thereby improving the OCR model's recognition ability for various types of images.
[0169] Based on the sample generation method provided in the above embodiment, this application also provides a sample generation device. Figure 8 Provide explanation. Figure 8 This is a schematic diagram of the structure of a sample generation device provided in an embodiment of the present application, the device comprising:
[0170] A scene construction module 801 is configured to construct a target virtual scene using a game engine; the target virtual scene includes model elements with text displayed on their surfaces;
[0171] An image rendering module 802 is configured to render the target virtual scene through the game engine to obtain a target scene image;
[0172] An information extraction module 803 is configured to run a text information extraction script during the rendering of the target virtual scene to extract text on the model elements in the target virtual scene and determine position information corresponding to the text in the target scene image;
[0173] The sample determination module 804 is configured to determine the label corresponding to the target scene image based on the extracted text and the corresponding position information; and determine that the target scene image and the corresponding label are training samples for an optical character recognition (OCR) model.
[0174] Optionally, the scene construction module 801 includes:
[0175] A creation unit, configured to create an initial virtual scene using the game engine;
[0176] A first deployment unit, configured to deploy three-dimensional model elements in the initial virtual scene;
[0177] The first selection unit is configured to select a target model element from the deployed three-dimensional model elements, and deploy a text object carrying text on a surface of the target model element to obtain the target virtual scene.
[0178] Optionally, the first deployment unit includes:
[0179] A second deployment unit is configured to run a scene layout script to deploy three-dimensional model elements in the initial virtual scene according to a preset model element deployment rule;
[0180] Correspondingly, the first selection unit includes:
[0181] The second selection unit is used to run the scene layout script to select the target model element from the deployed three-dimensional model elements according to the preset model selection rules; and to select the target text that is compatible with the target model element from the pre-built candidate text library according to the preset text deployment rules, and deploy a text object carrying the target text on the surface of the target model element.
[0182] Optionally, the text displayed on the surface of the model element is generated in the following manner:
[0183] a first determining unit configured to determine model reference information, wherein the model reference information includes at least one of prompt text and model reference parameters, wherein the prompt text is used to indicate content features corresponding to the text to be generated, and the model reference parameters are used to control basic features corresponding to the output results of the large language model;
[0184] The first generating unit is configured to call the large language model and generate the text based on the model reference information.
[0185] Optionally, the device further comprises:
[0186] an iterative unit, configured to, when the text generated by calling the large language model does not meet the preset text requirements, perform an information iterative adjustment operation until a text meeting the preset text requirements is obtained;
[0187] An adjustment unit, wherein the information iterative adjustment operation includes: adjusting the model reference information used when calling the large language model this time to obtain adjusted model reference information; calling the large language model again to generate text based on the adjusted model reference information.
[0188] Optionally, the image rendering module 802 includes:
[0189] A second determining unit is configured to determine rendering scene parameters; the rendering scene parameters are used to determine at least one of a lighting condition corresponding to the rendered scene image, a viewing angle corresponding to the scene image in the target virtual scene, and a material and a texture of a model element in the scene image;
[0190] A rendering unit is used to render the target virtual scene based on the rendering scene parameters through the game engine to obtain the target scene image.
[0191] Optionally, the device further comprises:
[0192] a first diffusion processing unit, configured to perform diffusion processing on the target scene image using a stable diffusion algorithm to obtain a diffused scene image; wherein the diffusion processing is configured to perform deformation processing on the entire target scene image or a local area in the target scene image;
[0193] The third determining unit is configured to determine a label corresponding to the diffuse scene image according to the label corresponding to the target scene image; and determine the diffuse scene image and the label corresponding thereto as training samples for the OCR model.
[0194] Optionally, the first diffusion processing unit includes:
[0195] a fourth determining unit, configured to determine a diffusion processing parameter; the diffusion processing parameter is used to indicate at least one of a diffusion step size, a number of diffusion iterations, a deformation intensity, a deformation effect, and a deformation region range;
[0196] The second diffusion processing unit is configured to perform diffusion processing on the target scene image based on the diffusion processing parameters by using the stable diffusion algorithm to obtain the diffused scene image.
[0197] Optionally, the information extraction module 803 includes:
[0198] The first acquisition unit is used to run the text information extraction script to obtain the text on the model element in the target virtual scene from the text component of the game engine; and to determine the positioning point corresponding to the display area of the text through the positioning component in the text component, convert the positioning point into corresponding screen coordinates, and obtain the position information corresponding to the text.
[0199] Optionally, the device further comprises:
[0200] A second acquisition unit is configured to acquire a training sample set, wherein the training sample set includes training samples generated based on a virtual scene constructed by the game engine;
[0201] a first recognition unit, configured to recognize the image in the training sample using the OCR model to be trained, and obtain predicted text information corresponding to the training sample; the predicted text information includes the predicted text included in the image and predicted position information corresponding to the predicted text;
[0202] The first training unit is configured to train the OCR model based on the labels of the training samples in the training sample set and the corresponding predicted text information.
[0203] Optionally, the training sample set further includes training samples generated based on real scenarios; and the apparatus further includes:
[0204] a classification unit, configured to classify the images in the training samples using the OCR model to obtain a predicted image type corresponding to the training samples; the predicted image type is used to indicate whether the image predicted by the OCR model is generated based on a virtual scene or a real scene;
[0205] Correspondingly, the first training unit includes:
[0206] A first construction unit is configured to construct a first loss function based on the labels of the training samples in the training sample set and the corresponding predicted text information;
[0207] A second construction unit is configured to construct a second loss function based on the predicted image type and the actual image type corresponding to each of the training samples in the training sample set;
[0208] The second training unit is used to train the OCR model according to the first loss function and the second loss function.
[0209] The embodiment of the present application also provides a computer device, which may specifically be a terminal device or a server. The terminal device and server provided in the embodiment of the present application will be introduced below from the perspective of hardware entity.
[0210] See also Figure 9 , Figure 9 This is a schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 9 For ease of explanation, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal can be any terminal device including a mobile phone, tablet computer, personal digital assistant (PDA), point of sales (POS), car computer, etc. For example, the terminal is a computer:
[0211] Figure 9 FIG2 is a block diagram showing a partial structure of a computer related to a terminal provided in an embodiment of the present application. Figure 9 The computer includes: a radio frequency (RF) circuit 1210, a memory 1220, an input unit 1230 (including a touch panel 1231 and other input devices 1232), a display unit 1240 (including a display panel 1241), a sensor 1250, an audio circuit 1260 (connected to a speaker 1261 and a microphone 1262), a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290. Those skilled in the art will understand that Figure 9 The computer structure shown in the figure does not constitute a limitation of the computer, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0212] Memory 1220 can be used to store software programs and modules. Processor 1280 executes the various computer functions and data processing by running the software programs and modules stored in memory 1220. Memory 1220 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function). The data storage area may store data generated based on the use of the computer (such as audio data, a phone book, etc.). Memory 1220 may also include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state memory device.
[0213] Processor 1280 is the computer's control center, connecting various computer components using various interfaces and circuits. It executes software programs and / or modules stored in memory 1220 and accesses data stored in memory 1220 to perform various computer functions and process data. Optionally, processor 1280 may include one or more processing units. Preferably, processor 1280 integrates an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1280.
[0214] In an embodiment of the present application, the processor 1280 included in the terminal is used to execute the steps in the methods described in the aforementioned embodiments.
[0215] See also Figure 10 , Figure 10 A structural diagram of a server 1300 provided for an embodiment of the present application. The server 1300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1322 (for example, one or more processors) and a memory 1332, and one or more storage media 1330 (for example, one or more massive storage devices) for storing application programs 1342 or data 1344. Among them, the memory 1332 and the storage medium 1330 may be temporary storage or permanent storage. The program stored in the storage medium 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1322 may be configured to communicate with the storage medium 1330 to execute a series of instruction operations in the storage medium 1330 on the server 1300.
[0216] The server 1300 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and / or one or more operating systems, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.
[0217] The steps performed by the server in the above embodiment can be based on the Figure 10 The server structure shown.
[0218] The CPU 1322 is configured to execute the steps of the methods described in the aforementioned embodiments.
[0219] An embodiment of the present application further provides a computer-readable storage medium for storing a computer program, which is used to execute the steps in the methods described in the aforementioned embodiments.
[0220] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the methods described in the aforementioned embodiments.
[0221] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0222] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0223] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0224] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0225] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store computer programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0226] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0227] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A sample generation method, characterized in that: The method comprises: Constructing a target virtual scene using a game engine; the target virtual scene includes a model element with text displayed on the surface; Rendering the target virtual scene by the game engine to obtain a target scene image; During the rendering of the target virtual scene, running a text information extraction script to extract text on the model elements in the target virtual scene and determine position information corresponding to the text in the target scene image; According to the extracted text and the corresponding position information, a label corresponding to the target scene image is determined; and the target scene image and the corresponding label are determined to be training samples of an optical character recognition (OCR) model.
2. The method according to claim 1, characterized in that The target virtual scene is constructed by using a game engine, including: Creating an initial virtual scene through the game engine; deploying three-dimensional model elements in the initial virtual scene; A target model element is selected from the deployed three-dimensional model elements, and a text object carrying text is deployed on the surface of the target model element to obtain the target virtual scene.
3. The method according to claim 2, characterized in that The deploying of three-dimensional model elements in the initial virtual scene includes: Running a scene layout script to deploy three-dimensional model elements in the initial virtual scene according to preset model element deployment rules; The step of selecting a target model element from the deployed three-dimensional model elements and deploying a text object carrying text on the surface of the target model element to obtain the target virtual scene includes: Run the scene layout script to select the target model element from the deployed three-dimensional model elements according to the preset model selection rules; and select the target text that is compatible with the target model element from the pre-built candidate text library according to the preset text deployment rules, and deploy a text object carrying the target text on the surface of the target model element.
4. The method according to any one of claims 1 to 3, characterized in that The text displayed on the surface of the model element is generated in the following way: Determining model reference information; the model reference information includes at least one of prompt text and model reference parameters, the prompt text is used to indicate content features corresponding to the text to be generated, and the model reference parameters are used to control basic features corresponding to the output results of the large language model; The large language model is called to generate the text based on the model reference information.
5. The method according to claim 4, characterized in that The method further comprises: If the text generated by calling the large language model does not meet the preset text requirements, performing an information iterative adjustment operation until a text meeting the preset text requirements is obtained; The information iteration adjustment operation includes: adjusting the model reference information used when calling the large language model this time to obtain adjusted model reference information; calling the large language model again to generate text based on the adjusted model reference information.
6. The method according to any one of claims 1 to 5, characterized in that The step of rendering the target virtual scene by the game engine to obtain a target scene image includes: Determining rendering scene parameters; the rendering scene parameters are used to determine at least one of the lighting conditions corresponding to the rendered scene image, the viewing angle corresponding to the scene image in the target virtual scene, and the material and texture of the model elements in the scene image; The target virtual scene is rendered by the game engine based on the rendering scene parameters to obtain the target scene image.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Using a stable diffusion algorithm, the target scene image is subjected to diffusion processing to obtain a diffused scene image; the diffusion processing is used to perform deformation processing on the entire target scene image or a local area in the target scene image; According to the label corresponding to the target scene image, the label corresponding to the diffuse scene image is determined; and the diffuse scene image and the label corresponding thereto are determined as training samples of the OCR model.
8. The method according to claim 7, characterized in that The method of adopting a stable diffusion algorithm to perform diffusion processing on the target scene image to obtain a diffused scene image includes: Determining diffusion processing parameters; the diffusion processing parameters are used to indicate at least one of a diffusion step size, a number of diffusion iterations, a deformation intensity, a deformation effect, and a deformation region range; The stable diffusion algorithm is used to perform diffusion processing on the target scene image based on the diffusion processing parameters to obtain the diffused scene image.
9. The method according to any one of claims 1 to 8, characterized in that The step of running the text information extraction script to extract the text on the model element in the target virtual scene and determining the position information corresponding to the text in the target scene image includes: Run the text information extraction script to obtain the text on the model element in the target virtual scene from the text component of the game engine; and determine the positioning point corresponding to the display area of the text through the positioning component in the text component, convert the positioning point into corresponding screen coordinates, and obtain the position information corresponding to the text.
10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: Acquire a training sample set; the training sample set includes training samples generated based on a virtual scene constructed by the game engine; Recognizing the image in the training sample by the OCR model to be trained to obtain predicted text information corresponding to the training sample; the predicted text information includes the predicted text included in the image and the predicted position information corresponding to the predicted text; The OCR model is trained based on the labels included in each of the training samples in the training sample set and the corresponding predicted text information.
11. The method according to claim 10, characterized in that The training sample set also includes training samples generated based on real scenes; the method further includes: Classifying the images in the training samples using the OCR model to obtain predicted image types corresponding to the training samples; the predicted image types are used to characterize whether the images predicted by the OCR model are generated based on a virtual scene or a real scene; The step of training the OCR model based on the labels of the training samples in the training sample set and the corresponding predicted text information includes: Constructing a first loss function based on the labels of the training samples in the training sample set and the corresponding predicted text information; Constructing a second loss function based on the predicted image type and the actual image type corresponding to each of the training samples in the training sample set; The OCR model is trained according to the first loss function and the second loss function.
12. A sample generating device, characterized in that: The device comprises: A scene construction module is used to construct a target virtual scene through a game engine; the target virtual scene includes a model element with text displayed on the surface; An image rendering module, configured to render the target virtual scene through the game engine to obtain a target scene image; An information extraction module is used to run a text information extraction script during the rendering of the target virtual scene to extract the text on the model elements in the target virtual scene and determine the position information corresponding to the text in the target scene image; The sample determination module is used to determine the label corresponding to the target scene image based on the extracted text and the corresponding position information; and determine that the target scene image and the corresponding label are training samples of the optical character recognition (OCR) model.
13. A computer device, characterized in that: The device includes a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the sample generation method according to any one of claims 1 to 11 according to the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is executed by an electronic device, the sample generation method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the sample generation method according to any one of claims 1 to 11 is implemented.