Method, device, electronic device, storage medium and computer program for identifying the meaning of a sign

The method and apparatus enhance sign recognition and interpretation in images by using pre-trained models, addressing the lack of efficient sign understanding in current technologies and improving user experience through accurate and comprehensive data presentation.

JP7813322B2Active Publication Date: 2026-02-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024157260
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-05-13
Filing Date
2024-09-11
Publication Date
2026-02-12
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

Current products do not support efficient recognition and understanding of the meaning of signs in images, such as laundry instructions and traffic signs, which are commonly encountered in daily life.

Method used

A method and apparatus for identifying the meaning of signs in images using pre-trained sign object detection and meaning identification models, capable of recognizing and interpreting signs in target images, including multiple sub-sign objects and considering contextual information like geographical location and user beliefs.

Benefits of technology

Improves the efficiency of acquiring and understanding sign data by automatically recognizing and interpreting signs, enhancing user information acquisition and providing comprehensive data presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813322000001
    Figure 0007813322000001
  • Figure 0007813322000002
    Figure 0007813322000002
  • Figure 0007813322000003
    Figure 0007813322000003
Patent Text Reader

Abstract

To provide a method, an apparatus, an electronic device, a storage medium, and a computer program product for specifying the meaning of a mark.SOLUTION: The method includes steps of determining whether or not a mark object is included in a target image (201), and specifying a meaning represented by the mark object according to the determination that the mark object is included in the target image (202).SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of computer technology, specifically to the field of image recognition and artificial intelligence, and in particular to a method and apparatus, electronic device, storage medium and computer program for identifying the meaning of a sign, which can be applied to the scene of sign meaning recognition. [Background technology]

[0002] In our daily lives, we see various signs, such as laundry instructions and traffic signs. When we come across a sign we don't know, we want to know its meaning. However, currently available products do not support sign recognition. Summary of the Invention [Means for solving the problem]

[0003] The present disclosure provides a method, an apparatus, an electronic device, a storage medium, and a computer program for determining the meaning of a sign.

[0004] According to a first aspect, there is provided a method for identifying the meaning of a sign, the method including the steps of determining whether a target image contains a sign object, and identifying the meaning represented by the sign object in response to determining that the target image contains a sign object.

[0005] According to a second aspect, there is provided an apparatus for identifying the meaning of a sign, the apparatus comprising: a sign identification unit configured to determine whether a target image includes a sign object; and a meaning identification unit configured to identify the meaning represented by the sign object in response to determining that the target image includes the sign object.

[0006] According to a third aspect, there is provided an electronic device comprising at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform a method according to any of the embodiments of the first aspect.

[0007] According to a fourth aspect, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions causing a computer to perform a method according to any embodiment of the first aspect.

[0008] According to a fifth aspect, there is provided a computer program which, when executed by a processor, implements the method according to any embodiment of the first aspect.

[0009] According to the technology of the present disclosure, a method and device for identifying the meaning of a sign are provided, which can automatically recognize whether a target image contains a sign object, and if the target image contains a sign object, can automatically identify the meaning represented by the sign object.By using signs efficiently expressed by a user through the target image, the efficiency of acquiring sign data and the efficiency of identifying the sign meaning are improved.

[0010] It should be noted that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor do they limit the scope of the present disclosure. Other features of the present disclosure will be easier to understand from the following description. The drawings are used for a better understanding of the present disclosure and are not a limitation to the present disclosure. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 illustrates an exemplary system architecture to which an embodiment of the present disclosure can be applied. [Figure 2]1 is a flow chart illustrating one embodiment of a method for determining the meaning of a sign according to the present disclosure. [Figure 3] 10A and 10B are schematic diagrams illustrating a plurality of types of sub-label objects in a washing indication. [Figure 4] FIG. 1 is a schematic diagram illustrating an application scenario of a method for identifying the meaning of a sign according to an embodiment of the present disclosure. [Figure 5] 10 is a flow chart illustrating another embodiment of a method for determining the meaning of a sign according to the present disclosure. [Figure 6] 1 is a structural diagram of one embodiment of an apparatus for determining the meaning of a sign according to the present disclosure; [Figure 7] FIG. 1 is a structural schematic diagram of a computer system suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are described below to facilitate understanding. However, it should be understood that these details are merely illustrative. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments of the present disclosure without departing from the scope and spirit of the present disclosure. In the following description, for clarity and simplicity, well-known functions and structures will not be described.

[0013] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and other processes of relevant user personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and morals.

[0014] FIG. 1 illustrates an exemplary system architecture 100 to which the method and apparatus for determining the meaning of signs according to the present disclosure may be applied.

[0015] 1, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Terminal devices 101, 102, and 103 are communicatively connected to form a topology network, and network 104 is used as a medium to provide a communication link between terminal devices 101, 102, and 103 and server 105. Network 104 may include various types of connections, such as wired, wireless communication links, or fiber optic cables.

[0016] The terminal devices 101, 102, and 103 may be hardware devices or software that support network connections for data exchange and data processing. If the terminal devices 101, 102, and 103 are hardware, they may be various electronic devices that support functions such as network connection, information acquisition, interaction, display, and processing, including, but not limited to, smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. If the terminal devices 101, 102, and 103 are software, they may be installed in the electronic devices exemplified above. For example, they may be implemented as multiple software programs or software modules for providing distributed services, or as a single software program or software module. No particular limitation is imposed here.

[0017] The server 105 may be a server that provides various services, such as a back-processing server that acquires target images uploaded by users via the terminal devices 101, 102, and 103 and identifies the meanings of sign objects included in the target images. For example, the server 105 may be a cloud server.

[0018] The server may be hardware or software. If the server is hardware, it may be implemented as a distributed server cluster consisting of multiple servers, or as a single server. If the server is software, it may be implemented as multiple pieces of software or software modules (e.g., software or software modules for providing distributed services), or as a single piece of software or software module. There are no particular limitations here.

[0019] The method for identifying the meaning of a sign according to the embodiment of the present disclosure may be executed by a server, a terminal device, or a server and a terminal device working together. Correspondingly, each part (e.g., each unit) included in the device for identifying the meaning of a sign may be entirely provided in the server, entirely provided in the terminal device, or even provided in both the server and the terminal device.

[0020] It should be understood that the number of terminal devices, networks, and servers in Figure 1 is merely exemplary. The number of terminal devices, networks, and servers may be arbitrarily increased or decreased according to the needs of implementation. If the electronic device on which the method for identifying the meaning of a sign operates does not need to perform data transmission with other electronic devices, the system architecture may include only the electronic device on which the method for identifying the meaning of a sign operates (e.g., a terminal device or a server).

[0021] Please refer to Figure 2. Figure 2 is a flow chart of a method for determining the meaning of a sign according to an embodiment of the present disclosure. Flow 200 includes the following steps:

[0022] In step 201, it is determined whether or not the target image contains a sign object.

[0023] In this embodiment, the entity executing the method for identifying the meaning of a sign (e.g., the server shown in Figure 1) may obtain a target image remotely or locally via a wired or wireless network connection, and determine whether the target image contains a sign object.

[0024] The sign represented by the sign object may be a sign in the narrow sense or a sign in the broad sense. In the narrow sense, it is various signs commonly seen in daily life, such as washing instructions, traffic signs, danger signs, and power signs. In the broad sense, it is a sign such as a brand or a flower.

[0025] For example, the execution entity may determine whether a target image contains a sign object using a pre-trained sign object detection model, where the sign object detection model is used to perform a classification operation on the target image to distinguish between images containing a sign object and images not containing a sign object.

[0026] A label object detection model can be either a one-stage or two-stage target detection model. In a one-stage target detection model, features are first extracted from the input image using a convolutional neural network (CNN). These features are then transmitted to a prediction head to predict the target bounding box and its category confidence at each location. Finally, techniques such as non-maximum suppression (NMS) are used to merge and filter overlapping bounding boxes to obtain the final detection result. Examples of such models include the Single Shot Multibox Detector (SSD), You Only Look Once (YOLO), and RetinaNet.

[0027] In a two-stage target detection model, image features are first extracted using a convolutional neural network and a candidate region extraction network (Region Proposal Network, RPN) is used to generate a set of candidate regions. These candidate regions are then converted into fixed-size profiles, and a classifier and regressor are used to perform object classification and high-precision localization for each candidate region. Mask R-CNN (Mask Region-based Convolutional Neural Network) then adds an additional branch to predict the mask for each target instance. Finally, techniques such as non-maximum suppression (NMS) are used to merge overlapping bounding boxes to obtain the final detection result. Examples of such models include Faster R-CNN (Faster Region-based Convolutional Neural Network), R-CNN (Region-based Convolutional Neural Network), and Mask R-CNN.

[0028] In step 202, in response to determining that the target image contains a sign object, the meaning expressed by the sign object is identified.

[0029] In this embodiment, in response to determining that a marker object is included in the target image, the execution subject can identify the meaning expressed by the marker object.

[0030] As an example, in response to determining that a sign object is included in the target image, the executing entity may use the sign recognition model to identify the specific category of sign object that the sign object belongs to, i.e., identify target category information of the sign object, and further identify the meaning corresponding to the target category information in the sign meaning table as the meaning expressed by the sign object.

[0031] The sign meaning table contains various signs and their corresponding semantic information, and may be generated by data statistics and organization.

[0032] As another example, the execution entity may identify the meaning represented by the sign object by a pre-trained meaning identification model, where the meaning identification model is used to represent the correspondence between the target image and the meaning of the sign object in the target image.

[0033] In some optional implementations of this embodiment, the execution entity may execute step 202 as follows.

[0034] In the first step, in response to determining that a sign object is included in the target image, a category to which the sign object belongs is identified.

[0035] In this embodiment, the markings may be classified based on common classification criteria, or may be classified based on preset classification criteria according to actual situations, for example, based on common classification criteria, the marking objects may be divided into categories such as laundry warning signs, traffic signs, danger signs, and power signs.

[0036] For example, in response to determining that a target image contains a marker object, the executing entity may perform feature extraction on the target image to obtain feature data, calculate the similarity between the feature data and each category feature in the category feature set, and identify the category corresponding to the category feature with the highest similarity as the category to which the marker object belongs, where each category feature in the category feature set represents one category.

[0037] In the second step, multiple types of sub-sign objects that belong to a category and are included in the target image are identified.

[0038] The target image often includes multiple sub-sign objects of different types in the same category. Next, refer to Figure 3. Figure 3 shows multiple sub-sign objects that belong to the category of washing instructions, such as "Do not bleach" and "Do not wash."

[0039] As an example, the executing entity divides the target image into sub-signal object types to obtain images of the sub-signal objects, performs feature extraction on the images of each sub-signal object to obtain sub-signal feature data, calculates the similarity between each sub-signal feature data and each sub-signal object reference feature in the sub-signal feature set corresponding to the category, and identifies the sub-signal object to which the sub-signal object reference feature corresponding to the greatest similarity belongs as one of the multiple types of sub-signal objects included in the target image.

[0040] Here, the sub-signal feature set corresponding to a category includes sub-signal object reference features of various sub-signal objects belonging to the category.

[0041] As another example, the executing entity may input a target image into a sign recognition model to identify multiple types of sub-sign objects belonging to a category that are included in the target image, where the sign recognition model is, for example, the one-stage target detection model or the two-stage target detection model.

[0042] In the third step, the meaning expressed by each of the multiple types of sub-sign objects is identified.

[0043] For each of the multiple types of sub-sign object, the executing entity may identify the meaning of the sub-sign object of that type by combining the sign recognition model and the sign meaning table, or may identify the meaning represented by the sub-sign object of that type by using the meaning identification model.

[0044] In this embodiment, the meanings of multiple sub-sign object types of different types in the same category can be simultaneously identified, improving the accuracy of the sign meaning identification process and the user's information acquisition efficiency.

[0045] In some optional implementations of this embodiment, the executing entity may perform the third step to identify, for each type of sub-marker object among multiple types of sub-marker objects, the meaning represented by that type of sub-marker object based on the number of sub-marker objects of that type in the target image.

[0046] For example, if the sign object is a flower, the meaning of the flower (floral language) varies depending on the type and number of flowers.

[0047] In this embodiment, for each of the multiple types of sub-marker objects, the executing entity first identifies the number of sub-marker objects of that type in the target image, and then identifies the meaning represented by the sub-marker object based on the category and number corresponding to the sub-marker object.

[0048] In this embodiment, the accuracy of the identified meaning is improved by identifying the meaning represented by the sub-sign object based on the category and number corresponding to the sub-sign object.

[0049] Continuing with the example of a sign object being a flower, the same type of flower may have different meanings depending on the geographical region, religious beliefs, and background. To further improve the accuracy of the identified meaning, the executing entity may acquire information on the user's geographical location, religious beliefs, background information, etc., and integrate the information to identify the meanings represented by various sub-sign object.

[0050] Next, reference is made to Fig. 4. Fig. 4 is a schematic diagram 400 showing an application scene of the method for identifying the meaning of a sign according to this embodiment. In the application scene of Fig. 4, a user takes a picture of a laundry care label on a terminal device 401 to obtain a target image 402, and uploads the target image 402 to a server 403. The server 403 first determines whether the target image includes a sign object, and then uploads the target image 402 to a server 403. sign In response to determining that the object is included, the meaning expressed by the sign object is identified and the meaning is returned to the terminal device.

[0051] In this embodiment, a method for identifying the meaning of a sign is provided, which automatically recognizes whether a sign object is included in a target image, and if the target image includes a sign object, it can automatically identify the meaning represented by the sign object. The sign efficiently expressed by the user through the target image improves the efficiency of acquiring sign data and identifying the sign meaning.

[0052] In some optional implementations of this embodiment, the execution entity may execute an operation to present multiple types of sub-sign objects and multiple meanings that correspond one-to-one to the multiple types of sub-sign objects.

[0053] In this embodiment, the execution subject may present a plurality of types of sub-marker objects and a plurality of meanings that correspond one-to-one to the plurality of types of sub-marker objects, using a preset presentation method.

[0054] Here, the preset presentation method may be set according to the actual situation, for example, by presenting a plurality of sign objects in one-to-one correspondence with a plurality of meanings.

[0055] In this embodiment, based on the multiple types of sub-signal objects and multiple meanings presented, the user can quickly identify the meanings of various sub-signal objects, thereby improving the user's information acquisition efficiency.

[0056] In some optional implementations of this embodiment, the execution entity may execute the presentation process as follows.

[0057] In the first step, images of a plurality of sub-sign objects that correspond one-to-one to a plurality of types of sub-sign objects are recalled.

[0058] In one example, the execution entity or an electronic device communicatively connected to the execution entity is provided with a sign image set, which includes reference images of various signs.

[0059] The execution entity can recall images of a plurality of sub-sign object that correspond one-to-one to the plurality of types of sub-sign object from the sign image set.

[0060] In the second step, images of multiple sub-sign objects and their meanings are presented in the form of a picture-text correspondence.

[0061] In this embodiment, the "figure" in the "figure-text correspondence format" refers to the image of the sub-sign object, and the "text" refers to its meaning. By adopting the figure-text correspondence format, the presentation effect of multiple sub-sign object images and multiple meanings is improved, contributing to improving the user's information acquisition efficiency.

[0062] In some optional implementations of this embodiment, the execution entity may execute the first step as follows.

[0063] First, for each of the plurality of types of sub-marker objects, in response to determining that a preset marker image set contains an image of a sub-marker object corresponding to that type of sub-marker object, an image of the sub-marker object is recalled from the preset marker image set.

[0064] As an example, for each of the multiple types of sub-signpost objects, the executing entity calculates the similarity between the feature data of the sub-signpost object of that type and the feature data of the image of each sub-signpost object in the preset signpost image set, and if the maximum similarity is greater than the preset similarity threshold, identifies and recalls the image of the sub-signpost object corresponding to the maximum similarity as the image of the sub-signpost object corresponding to that type of sub-signpost object, and if the maximum similarity is equal to or less than the preset similarity threshold, determines that the preset signpost image set does not include an image of a sub-signpost object corresponding to that type of sub-signpost object.

[0065] Then, in response to determining that the image of the sub-signal object corresponding to the sub-signal object of the type is not included in the preset signal image set, an image of the sub-signal object corresponding to the sub-signal object of the type is recalled from the image resources of the network.

[0066] The network image data is large in scale and has abundant resources. When it is determined that the predetermined sign image set does not include an image of a sub sign object corresponding to the type of sub sign object, an image of a sub sign object corresponding to the type of sub sign object may be recalled from the network image resources by similarity-based recall.

[0067] In this embodiment, a method for recalling images of sub-signal objects is provided, and the accuracy of the identified images of sub-signal objects is improved by combining a pre-set set of sign images with the image resources of the network.

[0068] In some optional implementations of this embodiment, the executing entity may perform the following operation: generate and present summary data for sign objects based on multiple types of sub-sign objects and multiple meanings that correspond one-to-one to the multiple types of sub-sign objects using a pre-trained artificial intelligence large-scale model.

[0069] Large-scale models (artificial intelligence large-scale models) specifically refer to pre-trained large-scale models of artificial intelligence with large or ultra-large parameters. Large-scale models of artificial intelligence have two meanings: "pre-training" and "large-scale model." Combining these two terms creates a new artificial intelligence mode, namely, a model that is pre-trained on a large dataset and can support various applications with only a small amount of fine-tuning or no fine-tuning. Large-scale models of artificial intelligence have excellent context understanding, language generation, learning, and transferability.

[0070] Taking the case where the sign object is a laundry symbol as an example, the executing entity can use the large-scale artificial intelligence model to generate comprehensive data to guide the user on how to wash laundry based on multiple types of sub-sign objects and multiple meanings that correspond one-to-one to the multiple types of sub-sign objects.

[0071] Taking the sign object as an example, when the sign object is a traffic sign, the executing entity can use the artificial intelligence large-scale model to generate comprehensive data for guiding the user through the process of driving within the area corresponding to the sign object based on multiple types of sub-sign objects and multiple meanings that correspond one-to-one to the multiple types of sub-sign objects.

[0072] In this embodiment, the meanings of various sub-sign objects are presented to the user, and summary data is also presented, which not only resolves the user's main question of "What do the signs mean?" but also satisfies the user's secondary request of "presenting a prompt," further improving the user's information acquisition efficiency and experience.

[0073] In some optional implementations of this embodiment, the executing entity may execute the process of generating the summary data as follows.

[0074] First, the prompt of the artificial intelligence large-scale model is determined according to the scene represented by the sign object.

[0075] In this embodiment, for each scene, a prompt corresponding to the scene is displayed. mp t) may be set in advance. For example, for a laundry label, the prompt information is "Please tell me how to care for the laundry according to the xxx on the laundry label. Just summarize it in one sentence," and for a flower sign, the prompt information is "Please tell me the number, name, and meaning of the corresponding flower."

[0076] Once the scene represented by the sign object is identified, the prompts of the artificial intelligence large scale model in that scene can be determined.

[0077] Then, an artificial intelligence large-scale model generates and presents summary data based on the prompt, multiple types of sub-sign objects, and multiple meanings.

[0078] The summary data generated as an example of a prompt corresponding to the above laundry label is, for example: Clothing care advice: follow normal washing instructions, hand wash in water below 30°C, do not bleach or tumble dry, iron at a temperature below 110°C, and dry cleaning is not recommended.

[0079] In this embodiment, by setting different prompts for different scenes, the large-scale model can generate and present summary data based on multiple types of sub-sign objects and multiple meanings, further improving the suitability of the generated summary data to the scene.

[0080] In some optional implementations of this embodiment, the execution entity may perform the following operations:

[0081] First, in response to the determination that the target image does not include a marker object, other objects in the target image are identified.

[0082] Other objects are Texts to be translated in the field of translation, People, vehicles or other objects to be monitored in the field of intelligent surveillance and security; Pedestrians, vehicles, or obstacles in the field of autonomous driving, Lesions, tumors, and other lesions in the field of medical image analysis. Vehicles, people, buildings, etc. in the drone photography field, In the field of agricultural and environmental observation, we monitor the growth status of crops, areas affected by pests and diseases, etc. Including, but not limited to: Presents other objects and processed data for other objects in a two-column stream.

[0083] For the other objects, processing may be performed on the target image using the processing method corresponding to the other objects to obtain processed data, and the processed data for the other objects and the other objects may be presented in a two-column stream.

[0084] In this embodiment, when the target image does not contain a sign object, other objects in the target image are identified and processed data for the other objects is presented, thereby improving the comprehensiveness of processing for the target image.

[0085] Reference is now made to Figure 5, which shows a schematic flow 500 of a further embodiment of a method for identifying the meaning of a sign according to the present disclosure, which includes the following steps:

[0086] In step 501, it is determined whether or not the target image contains a sign object.

[0087] In step 502, in response to determining that the target image contains a sign object, the category to which the sign object belongs is identified.

[0088] In step 503, multiple types of sub-sign objects that belong to a category and are included in the target image are identified.

[0089] In step 504, the meaning expressed by each of the plurality of types of sub-sign object is identified.

[0090] In step 505, for each of the multiple types of sub-marker objects, an image of the sub-marker object is recalled from the preset marker image set in response to determining that the preset marker image set contains an image of the sub-marker object corresponding to that type of sub-marker object.

[0091] In step 506, in response to determining that the pre-set sign image set does not include an image of a sub-sign object corresponding to the type of sub-sign object, an image of a sub-sign object corresponding to the type of sub-sign object is recalled from an image resource on the network.

[0092] In step 507, images of a plurality of sub-sign objects and a plurality of meanings are presented in a figure-text correspondence format.

[0093] In step 508, a prompt for the artificial intelligence large scale model is determined according to the scene represented by the sign object.

[0094] In step 509, the artificial intelligence large-scale model generates and presents summary data based on the prompt, the multiple sub-sign objects, and the multiple meanings.

[0095] As can be seen from this embodiment, compared with the embodiment corresponding to FIG. 2, the flow 500 of the method for identifying the meaning of signs in this embodiment specifically describes the process of identifying the meaning of signs and the process of generating summary data, which further improves the efficiency and experience of users in obtaining information.

[0096] Further referring to FIG. 6, as an embodiment of the method shown in each of the above figures, the present disclosure provides an embodiment of an apparatus for identifying the meaning of a sign, which embodiment of the apparatus corresponds to the embodiment of the method shown in FIG. 2, and which can be specifically applied to various electronic devices.

[0097] As shown in FIG. 6 , the apparatus 600 for identifying the meaning of a sign comprises a sign identification unit 601 configured to determine whether a target image includes a sign object, and a meaning identification unit 602 configured to identify the meaning represented by the sign object in response to determining that the target image includes the sign object.

[0098] In some optional implementations of this embodiment, the meaning identification unit 602 is further configured, in response to determining that the target image contains a sign object, to identify a category to which the sign object belongs, identify multiple types of sub-sign object that belong to the category and are contained in the target image, and identify meanings represented by each of the multiple types of sub-sign object.

[0099] In some optional implementations of this embodiment, the meaning identification unit 602 is further configured to identify, for each type of sub-signal object of the plurality of types, the meaning represented by the sub-signal object of that type based on the number of sub-signal objects of that type in the target image.

[0100] In some optional implementations of this embodiment, the device may further comprise a presentation unit (not shown) configured to present multiple types of sub-sign objects and multiple meanings in one-to-one correspondence with the multiple types of sub-sign objects.

[0101] In some optional implementations of this embodiment, the presentation unit is further configured to recall images of a plurality of sub-sign objects having one-to-one correspondence with the plurality of types of sub-sign objects, and present the images of the plurality of sub-sign objects and their meanings in a form of picture-text correspondence.

[0102] In some optional implementations of this embodiment, the presentation unit is further configured to, for each of the plurality of types of sub-signal objects, recall an image of the sub-signal object from the preset signal image set in response to determining that the preset signal image set includes an image of the sub-signal object corresponding to the type of sub-signal object, and to recall an image of the sub-signal object corresponding to the type of sub-signal object from an image resource of the network in response to determining that the preset signal image set does not include an image of the sub-signal object corresponding to the type of sub-signal object.

[0103] In some optional implementations of this embodiment, the above-mentioned apparatus further comprises a generation unit (not shown) configured to generate and present summary data for the sign object based on multiple types of sub-sign objects and multiple meanings that correspond one-to-one to the multiple types of sub-sign objects using a pre-trained artificial intelligence large-scale model.

[0104] In some optional implementations of this embodiment, the generation unit further determines a prompt for the artificial intelligence large-scale model based on the scene represented by the sign object, and generates and presents summary data by the artificial intelligence large-scale model based on the prompt, multiple types of sub-sign objects, and multiple meanings.

[0105] In some optional implementations of this embodiment, the above-mentioned device further comprises an object identification unit (not shown) configured to identify other objects in the target image in response to determining that the target image does not include the sign object, and a data presentation unit (not shown) configured to present the other objects and processed data for the other objects in a two-column stream.

[0106] In this embodiment, a device for identifying the meaning of a sign is provided, which automatically recognizes whether a sign object is included in a target image, and if a sign object is included in the target image, it can automatically identify the meaning represented by the sign object. The sign efficiently expressed by the user through the target image improves the efficiency of acquiring sign data and identifying the sign meaning.

[0107] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device comprising at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform a method for determining the meaning of a sign as described in any of the above embodiments.

[0108] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium having stored thereon computer instructions for causing a computer to perform the method for determining the meaning of a sign described in any of the above embodiments.

[0109] An embodiment of the present disclosure provides a computer program capable of implementing the method for determining the meaning of a sign according to any of the above embodiments when executed by a processor.

[0110] 7 illustrates a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device may represent various forms of numerical computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. Note that the components, their interconnections, and their functions illustrated herein are merely exemplary and are not intended to limit the embodiments of the present disclosure described and / or claimed herein.

[0111] 7, the electronic device 700 comprises a computing unit 701 that can perform various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 702 or loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 can further store various programs and data necessary for the operation of the device 700. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0112] In the electronic device 700, several components are connected to the I / O interface 705, including an input unit 706 such as a keyboard, a mouse, etc., an output unit 707 such as various types of displays, speakers, etc., a storage unit 708 such as a magnetic disk, an optical disk, etc., and a communication unit 709 such as a network plug-in, a modem, a wireless communication transceiver, etc. The communication unit 709 enables the electronic device 700 to exchange information or data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0113] The computing unit 701 may be various general-purpose and / or special-purpose processing components having processing and computational capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, and microcontroller. The computing unit 701 executes various methods and processes, such as the method for identifying the meaning of a sign described above. For example, in some embodiments, the method for identifying the meaning of a sign may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, it may perform one or more steps of the method for identifying the meaning of a sign described above. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the method for determining the meaning of the sign in any other suitable manner (eg, via firmware).

[0114] Various embodiments of the systems and techniques described herein may be realized in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments may be implemented in one or more computer programs that can be executed and / or interpreted in a programmable system that includes at least one programmable processor, which may be a special purpose or general purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0115] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a general-purpose computer, a special-purpose computer, or other processor or controller for an apparatus for identifying the meaning of signs, and when executed by the processor or controller, the functions or operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on the device, partially on the device, partially on the device as a stand-alone software package and partially on a remote device, or entirely on a remote device or server.

[0116] In the context of this disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use by or in connection with an instruction-execution system, device, or apparatus. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of machine-readable storage media may include a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0117] To provide for user interaction, the systems and techniques described herein can be implemented on a computer that includes a display device (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices can also be used to interact with a user. For example, the feedback provided to the user can be any form of sensing feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and can receive input from the user in any form, including sound input, speech input, or tactile input.

[0118] The systems and techniques described herein may be implemented in a computing system including back-end components (e.g., a data server), a computing system including middleware components (e.g., an application server), a computing system including front-end components (e.g., a user computer having a graphical user interface or a web browser) through which a user may interact with embodiments of the systems and techniques described herein, or a computing system including any combination of such back-end, middleware, or front-end components. The components of the system may be connected by digital data communication via any form or medium, such as a communications network, including a local area network (LAN), a wide area network (WAN), and the Internet.

[0119] The computer system may include a client and a server. The client and server are typically remote from each other and communicate with each other via a communication network. The relationship between the client and the server is created by running computer programs having a client-server relationship on each computer. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in a cloud computing service system and is configured to solve the drawbacks of traditional physical hosts and virtual private server (VPS) services, such as high management difficulty and poor business scalability, and may be a server in a distributed system or a server combined with a blockchain.

[0120] According to the technical solution of an embodiment of the present disclosure, a method for identifying the meaning of a sign is provided, which can automatically recognize whether a target image contains a sign object, and if the target image contains a sign object, can automatically identify the meaning represented by the sign object. The sign efficiently expressed by the user through the target image improves the efficiency of acquiring sign data and the efficiency of identifying the meaning of the sign.

[0121] It should be understood that steps can be rearranged, added, or deleted using the various forms of flow described above. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved. This specification does not limit the scope of the present disclosure.

[0122] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and principles of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. determining whether the target image includes a sign object; identifying a category to which the sign object belongs in response to determining that the target image contains the sign object; Identifying a plurality of types of sub-sign objects included in the target image and belonging to the category; A step of identifying the meaning represented by each of the plurality of types of sub-sign object; determining a prompt corresponding to the scene according to the scene represented by the sign object; generating and presenting summary data for the sign object based on the prompt, the plurality of types of sub-sign objects, and a plurality of meanings that correspond one-to-one to the plurality of types of sub-sign objects, using a pre-trained artificial intelligence large-scale model; A method for identifying the meaning of signs, including:

2. The step of identifying the meanings represented by each of the plurality of types of sub-sign object includes: The method according to claim 1 , further comprising a step of identifying, for each type of the plurality of types of sub-marker objects, a meaning expressed by the sub-marker objects of that type based on the number of sub-marker objects of that type in the target image.

3. The method of claim 1 , further comprising the step of presenting the plurality of types of sub-sign objects and a plurality of meanings that correspond one-to-one to the plurality of types of sub-sign objects.

4. The step of presenting the plurality of types of sub-sign object and the plurality of meanings corresponding one-to-one to the plurality of types of sub-sign object includes: Recalling images of a plurality of sub-signal objects that correspond one-to-one to the plurality of types of sub-signal objects; The method of claim 3 , further comprising: presenting the images of the plurality of sub-sign objects and the plurality of meanings in a pictorial-text format.

5. The step of recalling images of a plurality of sub-sign objects that correspond one-to-one to the plurality of types of sub-sign objects includes: a step of recalling an image of the sub sign object from the preset sign image set in response to determining that a preset sign image set includes an image of a sub sign object corresponding to the type of sub sign object for each of the plurality of types of sub sign objects; The method of claim 4, further comprising: in response to determining that the preset sign image set does not include an image of a sub-sign object corresponding to the type of sub-sign object, recalling an image of a sub-sign object corresponding to the type of sub-sign object from an image resource of the network.

6. identifying other objects in the target image in response to determining that the target image does not include the marker object; and presenting the other objects and the processed data for the other objects in a two-column form.

7. 1. A device for identifying the meaning of a sign, comprising: a sign identification unit configured to determine whether a target image includes a sign object; a meaning identification unit configured to identify a category to which the sign object belongs, identify a plurality of types of sub-sign object included in the target image and belonging to the category, and identify meanings expressed by each of the plurality of types of sub-sign object, in response to determining that the target image contains the sign object; a generating unit configured to determine a prompt corresponding to a scene represented by the sign object according to the scene, and generate and present summary data for the sign object based on the prompt, the plurality of sub-sign object types, and a plurality of meanings that correspond one-to-one to the plurality of sub-sign object types through a pre-trained artificial intelligence large-scale model; A device for determining the meaning of a sign comprising:

8. The semantic identification unit further comprises: The device according to claim 7 , configured to identify, for each type of the plurality of types of sub-marker objects, a meaning expressed by the sub-marker object of that type based on the number of sub-marker objects of that type in the target image.

9. The device according to claim 7 , further comprising a presentation unit configured to present the plurality of types of sub-sign object and a plurality of meanings having one-to-one correspondence with the plurality of types of sub-sign object.

10. The presentation unit further comprises: The device according to claim 9, configured to recall images of a plurality of sub-sign objects that correspond one-to-one to the plurality of types of sub-sign objects; and present the images of the plurality of sub-sign objects and the plurality of meanings in a pictorial-text correspondence format.

11. The presentation unit further comprises: The device according to claim 10, configured to: recall, for each type of the plurality of types of sub-signal objects, an image of the sub-signal object from the preset sign image set in response to determining that a sub-signal object image corresponding to the type of sub-signal object is included in the preset sign image set; and recall, from an image resource on a network, an image of the sub-signal object corresponding to the type of sub-signal object in response to determining that the preset sign image set does not include an image of the sub-signal object corresponding to the type of sub-signal object.

12. an object identification unit configured to identify other objects in the target image in response to determining that the target image does not include the marker object; The apparatus of claim 7 , further comprising: a data presentation unit configured to present the other objects and processed data for the other objects in a two-column form.

13. at least one processor; a memory communicatively coupled to the at least one processor; An electronic device comprising:

7. An electronic device, characterized in that the memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method of any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: A non-transitory computer-readable storage medium, the computer instructions causing a computer to perform the method of any one of claims 1 to 6.

15. A computer program product which, when executed by a processor, implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image recognition model training and image recognition methods, devices, and electronic equipment

    CN111476284B

  • Traffic sign recognition method and device and training method and device of neural network model

    CN111488770A

  • Sign recognition system, sign recognition method, and sign recognition program

    JP2014092936A

  • Method, system and computer program product for providing driving assistance

    US20210089796A1