Object recognition method and device, electronic equipment and computer readable storage medium

By obtaining user voice data and vehicle image features, combining image segmentation and voice data, accurately identifying target objects, the problem of inaccurate object recognition in multi-object environments is solved, and the accuracy and efficiency of human-computer interaction is improved.

CN120356178APending Publication Date: 2025-07-22GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510253845.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When multiple objects exist, existing object recognition methods cannot accurately identify objects described by users, resulting in low recognition accuracy and affecting the human-computer interactive experience.

Method used

By obtaining the user's voice data characteristics and the vehicle's initial forward-view image, the area segmentation is performed using the image segmentation angle, the target features are extracted, and the similarity value is determined based on the voice data characteristics, so as to accurately identify the target object.

Benefits of technology

It improves the accuracy of object recognition, ensures that the identified objects are highly consistent with the user description, and improves the fluency and efficiency of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356178A_ABST
    Figure CN120356178A_ABST
Patent Text Reader

Abstract

The invention provides an object recognition method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining voice data features of a user, an initial front-view image of a vehicle, and an image segmentation angle, carrying out the region segmentation of the initial front-view image according to the image segmentation angle, and obtaining a target front-view image, and extracting features of objects in the target front view image to obtain target features, determining a similarity value according to the target features of any object and the voice data features, and determining a target object according to the similarity value. According to the invention, region segmentation is carried out on the initial front-view image according to the image segmentation angle, the similarity value is further determined by combining the target features of the object and the voice data features, and the target object is determined according to the similarity value, so that the region where each object is located in the initial front-view image can be accurately determined; therefore, the object conforming to the position indicated by the user can be identified more accurately, and the accuracy of object identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of intelligent networking technology, and in particular, to an object recognition method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] With the improvement of the intelligence level of automobiles, users have higher requirements for the interaction experience with vehicles. Functions such as "ask what you can see in front of the car" enable users to perform reference matching for questions and answers in front of the vehicle through natural speech, which relies on object recognition of the vehicle.

[0003] Related object recognition methods usually rely on visual models. However, when there are multiple objects in an image, there are deviations between the objects recognized by the visual model and the objects described by the user. For example: There are multiple vehicles in the picture, and the visual model cannot accurately identify the object corresponding to the user's natural language description from multiple vehicles. This results in the fact that in actual applications, the recognized object does not match the object that the user wants to recognize, thereby leading to low accuracy of object recognition and poor human-computer interaction experience. Summary of the Invention

[0004] Embodiments of the present application provide an object recognition method, apparatus, electronic device, and computer-readable storage medium, aiming to improve the problem of low accuracy of object recognition existing in related technologies.

[0005] In a first aspect, embodiments of the present application provide an object recognition method, which includes:

[0006] Obtain the voice data feature of the user, the initial front view image of the vehicle, and the image segmentation angle;

[0007] Perform region segmentation on the initial front view image according to the image segmentation angle to obtain a target front view image;

[0008] Extract the feature of the object in the target front view image to obtain a target feature;

[0009] Determine a similarity value according to the target feature of any object and the voice data feature;

[0010] Determine a target object according to the similarity value.

[0011] Technical effects: By performing regional segmentation on the initial front view image according to the image segmentation angle, it is possible to accurately identify the area where each object is located in the initial front view image, and then more precisely identify the object that matches the user's designated position, thereby improving the accuracy of object recognition. In addition, by combining the target features of the object and the voice data features, a similarity value is determined. This similarity value is a quantitative representation of the matching degree between the target features of the object and the voice data features. Then, the target object is determined based on the similarity value, ensuring that the recognized object highly matches the user's voice description. During the human-computer interaction process, it can more accurately respond to the user's needs, thereby effectively enhancing the human-computer interaction experience and bringing a smoother and more efficient interaction feeling to the user.

[0012] Optionally, the obtaining of the voice data features of the user includes:

[0013] Obtain the voice data of the user;

[0014] Perform speech recognition on the voice data to obtain text data;

[0015] Extract the features from the text data to obtain voice data features.

[0016] Technical effects: By extracting the features corresponding to the user's voice data, the key information in the user's needs can be accurately located, which helps the vehicle answer the user's questions more pertinently. In addition, using a preset feature word library to extract features can quickly screen out useful information, reducing the burden on the system for a comprehensive and complex analysis of the entire text data, thereby accelerating the system's response speed to the user's needs.

[0017] Optionally, the obtaining of the image segmentation angle includes:

[0018] Obtain the front view wide angle and vehicle speed of the vehicle;

[0019] Determine the image segmentation angle according to the front view wide angle and the vehicle speed.

[0020] Technical effects: By obtaining the current front view wide angle and vehicle speed of the vehicle in real time, the timeliness of the data is ensured. The determined image segmentation angle can timely reflect the current actual driving state of the vehicle, and the real-time updated data can quickly adjust the image segmentation angle, ensuring the accuracy of the initial front view image segmentation. In addition, by combining two key factors, the front view wide angle and the vehicle speed, to determine the image segmentation angle, compared with considering a single factor, the initial front view image can be divided more accurately, which helps to improve the accuracy of object recognition.

[0021] Optionally, the extracting of the features of the object in the target front view image to obtain the target features includes:

[0022] Extract the preliminary features of the object in the target forward-looking image;

[0023] Extract the orientation features of the object in the target forward-looking image;

[0024] Fuse the preliminary features and the orientation features to obtain target features.

[0025] Technical effect: By fusing the preliminary features and the orientation features to obtain target features, it can provide more comprehensive information for object recognition of the vehicle. Such fused target features combine the basic attributes and spatial position information of the object, greatly improving the probability of accurately identifying the object in a complex vehicle driving environment, being able to more precisely identify the object referred to by the user, and reducing the situation of misrecognition.

[0026] Optionally, the preliminary features include color features and contour features, and the voice data features include the orientation features, the color features, and the contour features;

[0027] Determining the similarity value according to the target features and the voice data features of any object includes:

[0028] Determine the similarity value of the color features according to the color features in the target features and the color features in the voice data features;

[0029] Determine the similarity value of the contour features according to the contour features in the target features and the contour features in the voice data features;

[0030] Determine the similarity value of the orientation features according to the orientation features in the target features and the orientation features in the voice data features;

[0031] Determine the similarity value according to the similarity value of the color features, the similarity value of the contour features, and the similarity value of the orientation features.

[0032] Technical effect: By comprehensively considering the similarity of features in three different aspects of color, contour, and orientation, it can more comprehensively evaluate the matching degree between the target features and the voice data features. In practical applications, considering only one feature alone may lead to misjudgment, while considering multiple features comprehensively can improve the accuracy and reliability of the judgment.

[0033] Optionally, determining the target object according to the similarity value includes:

[0034] Sort the similarity values to obtain the sorted similarity values;

[0035] Determine the maximum similarity value from the sorted similarity values;

[0036] Determine the object corresponding to the maximum similarity value as the target object.

[0037] Technical effect: By sorting these similarity values, it is possible to distinguish the degree of relevance to the user's intention among many objects. Further, taking the object corresponding to the maximum similarity value as the target object means that the target features of this object match the user's speech data features most closely. This is like selecting the most matching one among many candidates, improving the accuracy of determining the target object and avoiding misjudgment or selecting an object that does not match the user's intention.

[0038] Optionally, after the step of determining the target object according to the similarity value, the method includes:

[0039] Obtain an image of the target object;

[0040] Input the image of the target object into a pre-generated visual large model to obtain the attribute data of the target object.

[0041] Technical effect: By inputting the image of the target object into a pre-generated visual large model, the attribute data of the target object can be directly obtained, avoiding the complex and time-consuming steps such as manual feature extraction and classification in the traditional method, and greatly improving the speed of obtaining attribute data. For example, when identifying a vehicle, there is no need to manually judge the brand characteristics of the vehicle one by one, and the visual large model can quickly output the brand of the vehicle.

[0042] In a second aspect, an object recognition device provided by an embodiment of the present application includes:

[0043] A data acquisition module, configured to acquire the speech data features of a user, the front view image of a vehicle, and the image segmentation angle;

[0044] An image segmentation module, configured to perform region segmentation on the initial front view image according to the image segmentation angle to obtain a target front view image;

[0045] A feature extraction module, configured to extract the features of the object in the target front view image to obtain target features;

[0046] A similarity calculation module, configured to determine a similarity value according to the target features of any object and the speech data features;

[0047] An object recognition module, configured to determine a target object according to the similarity value.

[0048] Optionally, the data acquisition module includes:

[0049] A speech data acquisition sub-module, configured to acquire the speech data of the user;

[0050] A speech recognition sub-module for performing speech recognition on the speech data to obtain text data;

[0051] A first feature extraction sub-module for extracting features from the text data to obtain speech data features.

[0052] Optionally, the data acquisition module includes:

[0053] An angle and vehicle speed acquisition sub-module for acquiring the front view wide angle and vehicle speed of the vehicle;

[0054] An image segmentation angle determination sub-module for determining the image segmentation angle according to the front view wide angle and the vehicle speed.

[0055] Optionally, the feature extraction module includes:

[0056] A second feature extraction sub-module for extracting preliminary features of the object in the target front view image;

[0057] A third feature extraction sub-module for extracting the orientation features of the object in the target front view image;

[0058] A feature fusion sub-module for fusing the preliminary features and the orientation features to obtain target features.

[0059] Optionally, the preliminary features include color features and contour features, and the speech data features include the orientation features, the color features, and the contour features;

[0060] The similarity calculation module includes:

[0061] A first similarity calculation sub-module for determining the similarity value of the color features according to the color features in the target features and the color features in the speech data features;

[0062] A second similarity calculation sub-module for determining the similarity value of the contour features according to the contour features in the target features and the contour features in the speech data features;

[0063] A third similarity calculation sub-module for determining the similarity value of the orientation features according to the orientation features in the target features and the orientation features in the speech data features;

[0064] A target similarity calculation sub-module for determining the similarity value according to the similarity value of the color features, the similarity value of the contour features, and the similarity value of the orientation features.

[0065] Optionally, the object recognition module includes:

[0066] A similarity sorting sub-module for sorting the similarity values to obtain the sorted similarity values;

[0067] A maximum similarity determination sub-module for determining the maximum similarity value from the sorted similarity values;

[0068] A target object determination sub-module for determining the object corresponding to the maximum similarity value as the target object.

[0069] Optionally, the device includes:

[0070] An object image acquisition module for acquiring an image of the target object;

[0071] An object image processing module for inputting the image of the target object into a pre-generated visual large model to obtain attribute data of the target object.

[0072] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; a memory for storing instructions executable by the processor, wherein the processor is configured to execute the instructions to implement the object recognition method as described in any one of the above.

[0073] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the object recognition method as described in any one of the above.

[0074] In an embodiment of the present application, the voice data features of the user, the initial front view image of the vehicle, and the image segmentation angle are obtained, the initial front view image is regionally segmented according to the image segmentation angle to obtain a target front view image, the features of the objects in the target front view image are extracted to obtain target features, the similarity value is determined according to the target features and voice data features of any object, and the target object is determined according to the similarity value. By regionally segmenting the initial front view image according to the image segmentation angle, the present application can accurately clarify the region where each object in the initial front view image is located, and further can more accurately identify the object that conforms to the user's designated position, thereby improving the accuracy of object recognition. In addition, by combining the target features and voice data features of the object, the present application determines the similarity value, which is a quantitative manifestation of the matching degree between the target features and voice data features of the object, and then determines the target object according to the similarity value, ensuring that the recognized object highly coincides with the user's voice description. During the human-computer interaction process, the user's needs can be more accurately responded to, thereby effectively improving the human-computer interaction experience and bringing a more smooth and efficient interaction feeling to the user.

[0075] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. Description of the Drawings

[0076] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0077] Figure 1 is one of the flowcharts of the object recognition method provided by an embodiment of this application;

[0078] Figure 2 is a schematic diagram of regionally segmenting an initial front view image according to the image segmentation angle provided by an embodiment of this application;

[0079] Figure 3 is a schematic diagram of extracting the features of an object in a target front view image to obtain target features provided by an embodiment of this application;

[0080] Figure 4 is the second flowchart of the object recognition method provided by an embodiment of this application;

[0081] Figure 5 is the structural diagram of the object recognition device provided by an embodiment of this application;

[0082] Figure 6 is the structural diagram of the electronic device provided by an embodiment of this application. Detailed Embodiments

[0083] In order to make the technical problems, technical solutions and beneficial effects solved by this application more clear, the following further details this application in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0084] An object recognition method provided by an embodiment of this application includes: obtaining the voice data features of a user, the initial front view image of a vehicle, and the image segmentation angle, regionally segmenting the initial front view image according to the image segmentation angle to obtain a target front view image, extracting the features of an object in the target front view image to obtain target features, determining a similarity value according to the target features and voice data features of any object, and determining a target object according to the similarity value.

[0085] Technical effect: By performing regional segmentation on the initial forward view image according to the image segmentation angle, it is possible to accurately identify the region where each object in the initial forward view image is located, and then more precisely identify the object that conforms to the user's designated position, thereby improving the accuracy of object recognition. In addition, by combining the target features of the object and the voice data features, a similarity value is determined. This similarity value is a quantitative representation of the matching degree between the target features of the object and the voice data features. Then, based on the similarity value, the target object is determined, ensuring that the recognized object highly matches the user's voice description. During the human-computer interaction process, it can more accurately respond to the user's needs, thereby effectively enhancing the human-computer interaction experience and bringing a more smooth and efficient interaction feeling to the user.

[0086] Embodiment 1

[0087] An embodiment of the present application provides an object recognition method. Please refer to Figure 1 , and includes the following steps:

[0088] Step 101, obtain the voice data features of the user, the initial forward view image of the vehicle, and the image segmentation angle.

[0089] In some embodiments of the present application, obtaining the voice data features of the user is a key information source for understanding the user's intention. At the same time, the initial forward view image of the vehicle and the image segmentation angle are obtained through an in-vehicle camera. Among them, the initial forward view image is the original image of the scene in front of the current vehicle, and the image segmentation angle is the basis for subsequent image region segmentation.

[0090] Step 102, perform regional segmentation on the initial forward view image according to the image segmentation angle to obtain the target forward view image.

[0091] In some embodiments of the present application, through a suitable image segmentation angle, the initial forward view image can be segmented into different regions. The segmented initial forward view image is the target forward view image for more accurate object recognition in the future. For example: as Figure 2 shown, through the image segmentation angle, the initial forward view image in front of the vehicle is segmented into the front, left, and right regions.

[0092] Step 103, extract the features of the objects in the target forward view image to obtain the target features.

[0093] Step 104, determine the similarity value according to the target features and voice data features of any object.

[0094] Step 105, determine the target object according to the similarity value.

[0095] In some embodiments of the present application, a target detection algorithm can be used to detect objects in the target front view image, and the features of each object in the target front view image are extracted. The extracted features are the target features.

[0096] For any object in the target front view image, the similarity between the target features of the object and the speech data features is calculated to obtain a similarity value. If there are multiple objects in the target front view image, multiple similarity values are calculated. This step is a key step in associating visual information (features of the object) and speech information (speech data features).

[0097] After obtaining multiple similarity values, the multiple similarity values are compared to determine the target object that best matches the user's intention.

[0098] By regionally segmenting the initial front view image from the perspective of image segmentation, the present application can accurately identify the region where each object in the initial front view image is located, and thus can more precisely identify the object that matches the user's designated position, thereby improving the accuracy of object recognition. In addition, by combining the target features of the object and the speech data features, the present application determines a similarity value, which is a quantitative manifestation of the matching degree between the target features of the object and the speech data features. Then, based on the similarity value, the target object is determined, ensuring that the recognized object highly coincides with the user's speech description. During the human-computer interaction process, the user's needs can be more accurately responded to, thereby effectively enhancing the human-computer interaction experience and bringing a more smooth and efficient interaction feeling to the user.

[0099] Step 101 includes: obtaining the user's speech data, performing speech recognition on the speech data to obtain text data, and extracting the features in the text data to obtain speech data features.

[0100] In some embodiments of the present application, the user's speech data is obtained through a vehicle-mounted voice system. For example, the user's speech data can be "What brand is the white car on my left?".

[0101] After obtaining the user's speech data, the user's speech data is converted into corresponding text data. The features in the text data are extracted through a preset feature word library, and the extracted features are the speech data features. For example, if the preset feature word library includes feature words such as "orientation" and "color", then the speech data features are "left", "white", etc.

[0102] By extracting the features corresponding to the user's voice data, this application can accurately locate the key information in the user's needs, which helps the vehicle answer the user's questions more pertinently. In addition, this application extracts features using a preset feature word library, which can quickly screen out useful information, reduce the burden on the system for a comprehensive and complex analysis of the entire text data, and thus speed up the system's response speed to the user's needs.

[0103] Step 101 includes: obtaining the front view wide angle and vehicle speed of the vehicle, and determining the image segmentation angle according to the front view wide angle and vehicle speed.

[0104] In some embodiments of this application, the front view wide angle of the vehicle refers to the horizontal angle range that the front camera of the vehicle can cover.

[0105] Obtain the current front view wide angle and current vehicle speed of the vehicle in real time. Calculate the image segmentation angle in the current state of the vehicle according to the current front view wide angle and current vehicle speed of the vehicle.

[0106] Specifically, substituting the front view wide angle and vehicle speed of the vehicle into formula (1) can obtain the image segmentation angle.

[0107] F(W / 180) = Sin(At,Bg) Formula (1)

[0108] Wherein, F(W / 180) is a preset function with W / 180 as the independent variable, W is the image segmentation angle, A and B are both preset parameters, which can be adjusted subsequently with the optimization of the solution, t is the vehicle speed, and g is the front view wide angle.

[0109] By obtaining the current front view wide angle and vehicle speed of the vehicle in real time, this application ensures the timeliness of the data. The image segmentation angle determined in this way can reflect the current actual driving state of the vehicle in a timely manner, and the real-time updated data can quickly adjust the image segmentation angle, ensuring the accuracy of the initial front view image segmentation. In addition, by combining two key factors, the front view wide angle and vehicle speed, to determine the image segmentation angle, this application can divide the initial front view image more accurately compared to considering a single factor, which helps to improve the accuracy of object recognition.

[0110] Step 103 includes: extracting the preliminary features of the objects in the target front view image, extracting the orientation features of the objects in the target front view image, and fusing the preliminary features and orientation features to obtain the target features.

[0111] In some embodiments of this application, image processing algorithms are used to extract the preliminary features of the objects in the target front view image. These preliminary features are the basis for subsequent recognition and can help to initially distinguish different types of objects.

[0112] In the vehicle driving environment, the orientation feature of an object is very important for object recognition. Therefore, an image processing algorithm can be used to extract the orientation feature of an object in the front view image of the target, where the orientation feature includes the position of the object relative to the vehicle (e.g., directly in front of the vehicle, in the front left, in the front right, etc.).

[0113] Neither the individual preliminary features nor the orientation features alone are sufficient to accurately identify the object that meets the user's intention. For example: only knowing that the object is white and rectangular (preliminary features), but not knowing its orientation in front of the vehicle, it is impossible to accurately identify the object referred to by the user from multiple objects; conversely, only knowing that the object is directly in front of the vehicle (orientation feature), but not knowing whether it is a vehicle, a pedestrian or other object, it is also impossible to identify the object referred to by the user from multiple objects.

[0114] Therefore, the preliminary features and the orientation features can be fused, and the resulting feature after fusion is the target feature, which can provide more comprehensive information for object recognition of the vehicle.

[0115] Since neither the individual preliminary features nor the orientation features alone are sufficient to accurately identify the object that meets the user's intention, in this application, by fusing the preliminary features and the orientation features to obtain the target feature, it can provide more comprehensive information for object recognition of the vehicle. This fused target feature combines the basic attributes and spatial position information of the object, greatly improving the probability of accurately identifying the object in a complex vehicle driving environment, being able to more precisely identify the object referred to by the user, and reducing the situation of misidentification.

[0116] Step 104 includes: determining the similarity value of the color feature according to the color feature in the target feature and the color feature in the voice data feature, determining the similarity value of the contour feature according to the contour feature in the target feature and the contour feature in the voice data feature, determining the similarity value of the orientation feature according to the orientation feature in the target feature and the orientation feature in the voice data feature, and determining the similarity value according to the similarity value of the color feature, the similarity value of the contour feature, and the similarity value of the orientation feature.

[0117] In some embodiments of this application, as Figure 3 shown, the target feature contains a description of the color of the object, that is, the color feature, such as "blue", "white", "gray", "red", etc. At the same time, the voice data feature also contains a description of the color, that is, the color feature. Calculate the similarity between the color feature in the target feature and the color feature in the voice data feature to obtain a similarity value regarding the color feature.

[0118] The contour feature in the target feature describes the shape edge information of the object in front of the vehicle, such as: a vehicle, etc. The contour feature in the voice data feature may describe the object through voice, such as: a vehicle, etc. Calculate the similarity between the contour feature in the target feature and the contour feature in the voice data feature to obtain a similarity value regarding the contour feature.

[0119] The orientation feature in the target feature specifies the position information of the object in front of the vehicle relative to the vehicle, such as: left, front, right. The orientation feature in the voice data feature is the object orientation information described through voice. Calculate the similarity between the orientation feature in the target feature and the orientation feature in the voice data feature to obtain a similarity value regarding the orientation feature.

[0120] After obtaining the similarity value of the color feature, the similarity value of the contour feature, and the similarity value of the orientation feature, determine the similarity value according to the similarity value of the color feature, the similarity value of the contour feature, and the similarity value of the orientation feature. Specifically, the similarity value = the similarity value of the color feature + the similarity value of the contour feature + the similarity value of the orientation feature.

[0121] In this application, the final similarity value is obtained by adding the similarity value of the color feature, the similarity value of the contour feature, and the similarity value of the orientation feature. This calculation method is simple and intuitive, and can comprehensively consider the matching situations of multiple key features. Through this comprehensive calculation, an overall measurement index can be obtained to reflect the overall similarity between the target feature and the voice data feature. Further, because the similarity features of three different aspects of color, contour, and orientation are integrated, the matching degree between the target feature and the voice data feature can be evaluated more comprehensively. In practical applications, considering only a single feature may lead to misjudgment, while considering multiple features comprehensively can improve the accuracy and reliability of the judgment.

[0122] Step 105 includes: sorting the similarity values to obtain the sorted similarity values, determining the maximum similarity value from the sorted similarity values, and determining the object corresponding to the maximum similarity value as the target object.

[0123] In some embodiments of this application, since there may be multiple objects in front of the vehicle, there will be a corresponding similarity value between the target feature of each object and the voice data feature of the user. Sorting these similarity values can help the vehicle quickly determine which object best matches the user's intention from multiple objects, so as to more effectively screen out the target object.

[0124] Sorting can be carried out in descending or ascending order. Assuming that the similarity values are sorted in descending order, the similarity values ranked at the front after sorting are the maximum similarity values. The object corresponding to the maximum similarity value is more in line with the user's intention. Among them, the object corresponding to the maximum similarity value is the target object.

[0125] By sorting these similarity values in this application, it is possible to distinguish the degree of relevance to the user's intention among many objects. Further, taking the object corresponding to the maximum similarity value as the target object means that the target features of this object most conform to the user's speech data features. This is like selecting the most matching one among many candidates, improving the accuracy of determining the target object and avoiding misjudgment or selecting an object that does not match the user's intention.

[0126] Embodiment 2

[0127] An object recognition method is provided in an embodiment of this application. Please refer to Figure 4 , and after step 105, the following steps are further included:

[0128] Step 401, obtain an image of the target object.

[0129] Step 402, input the image of the target object into a pre-generated visual large model to obtain attribute data of the target object.

[0130] In some embodiments of this application, after the target object is recognized, an image of the target object can be obtained.

[0131] Taking the obtained image of the target object as the input of a pre-generated visual large model and inputting it into the visual large model, the visual large model can automatically output the attribute data of the target object. For example: if the target object is a vehicle, an image of the vehicle can be obtained. Taking the obtained image of the vehicle as the input of the visual large model and inputting it into the visual large model, the visual large model can automatically output the attribute data of the vehicle. Among them, the attribute data of the vehicle can include the brand, etc.

[0132] By inputting the image of the target object into a pre-generated visual large model in this application, the attribute data of the target object can be directly obtained, avoiding the complex and time-consuming steps such as manual feature extraction and classification in the traditional method, and greatly improving the speed of obtaining attribute data. For example: when recognizing a vehicle, there is no need to manually judge the brand features of the vehicle one by one, and the visual large model can quickly output the brand of the vehicle.

[0133] An object recognition device 80 is also provided in an embodiment of this application. Please refer to Figure 5 , including:

[0134] A data acquisition module 810, configured to acquire the voice data features of a user, the front view image of a vehicle, and the image segmentation angle;

[0135] An image segmentation module 820, configured to perform region segmentation on the initial front view image according to the image segmentation angle to obtain a target front view image;

[0136] A feature extraction module 830, configured to extract the features of an object in the target front view image to obtain target features;

[0137] A similarity calculation module 840, configured to determine a similarity value according to the target features of any object and the voice data features;

[0138] An object recognition module 850, configured to determine a target object according to the similarity value.

[0139] Optionally, the data acquisition module 810 includes:

[0140] A voice data acquisition sub-module, configured to acquire the voice data of a user;

[0141] A voice recognition sub-module, configured to perform voice recognition on the voice data to obtain text data;

[0142] A first feature extraction sub-module, configured to extract the features in the text data to obtain voice data features.

[0143] Optionally, the data acquisition module 810 includes:

[0144] An angle and vehicle speed acquisition sub-module, configured to acquire the front view wide angle and vehicle speed of a vehicle;

[0145] An image segmentation angle determination sub-module, configured to determine the image segmentation angle according to the front view wide angle and vehicle speed.

[0146] Optionally, the feature extraction module 830 includes:

[0147] A second feature extraction sub-module, configured to extract the preliminary features of an object in the target front view image;

[0148] A third feature extraction sub-module, configured to extract the orientation features of an object in the target front view image;

[0149] A feature fusion sub-module, configured to fuse the preliminary features and the orientation features to obtain target features.

[0150] Optionally, the preliminary features include color features and contour features, and the voice data features include orientation features, color features, and contour features;

[0151] The similarity calculation module 840 includes:

[0152] The first similarity operator module is used to determine the similarity value of the color feature according to the color feature in the target feature and the color feature in the voice data feature;

[0153] The second similarity operator module is used to determine the similarity value of the contour feature according to the contour feature in the target feature and the contour feature in the voice data feature;

[0154] The third similarity operator module is used to determine the similarity value of the orientation feature according to the orientation feature in the target feature and the orientation feature in the voice data feature;

[0155] The target similarity operator module is used to determine the similarity value according to the similarity value of the color feature, the similarity value of the contour feature, and the similarity value of the orientation feature.

[0156] Optionally, the object recognition module 850 includes:

[0157] The similarity sorting sub-module is used to sort the similarity values to obtain the sorted similarity values;

[0158] The maximum similarity determination sub-module is used to determine the maximum similarity value from the sorted similarity values;

[0159] The target object determination sub-module is used to determine the object corresponding to the maximum similarity value as the target object.

[0160] Optionally, the device includes:

[0161] The object image acquisition module is used to acquire the image of the target object;

[0162] The object image processing module is used to input the image of the target object into a pre-generated visual large model to obtain the attribute data of the target object.

[0163] This application embodiment also provides an electronic device 90, please refer to Figure 6 , which includes a processor 910 and a memory 920. Among them, the memory 910 is used to store a computer program; the processor 920 is used to execute the program stored on the memory 910 to implement the object recognition method introduced in any embodiment of this application.

[0164] This application embodiment also provides a computer-readable storage medium, which stores a computer program therein. When the computer program is executed by a processor, it implements the object recognition method introduced in any embodiment of this application.

[0165] In this application, multiple means two or more.

[0166] In this application, unless otherwise clearly defined, the terms "install", "connect", and "link" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0167] The terms "first", "second", "third", "fourth", etc. (if any) in this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0168] The term "and / or" in this application is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application generally indicates that the related objects before and after are in an "or" relationship.

[0169] If there is no special instruction, all steps of this application can be carried out in sequence or randomly. For example, the method includes steps A and B, which means that the method may include steps A and B carried out in sequence, or steps B and A carried out in sequence. For example, it is mentioned that the method may further include step C, which means that step C can be added to the method in any order. For example, the method may include steps A, B, and C, or steps A, C, and B, or steps C, A, and B, etc.

[0170] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. An object recognition method, characterized in that, The method includes: Obtaining the voice data features of the user, the initial front view image of the vehicle, and the image segmentation angle; Performing region segmentation on the initial front view image according to the image segmentation angle to obtain a target front view image; Extracting the features of the objects in the target front view image to obtain target features; Determining a similarity value according to the target features of any object and the voice data features; Determining the target object according to the similarity value.

2. The method according to claim 1, characterized in that, The obtaining of the voice data features of the user includes: Obtaining the voice data of the user; Performing speech recognition on the voice data to obtain text data; Extracting the features in the text data to obtain voice data features.

3. The method according to claim 1, wherein The obtaining of the image segmentation angle includes: Obtaining the front view wide angle and the vehicle speed of the vehicle; Determining the image segmentation angle according to the front view wide angle and the vehicle speed.

4. The method according to claim 1, wherein The extracting of the features of the objects in the target front view image to obtain target features includes: Extracting the preliminary features of the objects in the target front view image; Extracting the orientation features of the objects in the target front view image; Fusing the preliminary features and the orientation features to obtain target features.

5. The method according to claim 4, wherein The preliminary features include color features and contour features, and the voice data features include the orientation features, the color features, and the contour features; The determining of the similarity value according to the target features of any object and the voice data features includes: Determining the similarity value of the color features according to the color features in the target features and the color features in the voice data features; Determining the similarity value of the contour features according to the contour features in the target features and the contour features in the voice data features; Determining the similarity value of the orientation features according to the orientation features in the target features and the orientation features in the voice data features; Determining the similarity value according to the similarity value of the color features, the similarity value of the contour features, and the similarity value of the orientation features.

6. The method according to claim 1, wherein The determining of the target object according to the similarity value includes: Sorting the similarity values to obtain the sorted similarity values; Determining the maximum similarity value from the sorted similarity values; Determining the object corresponding to the maximum similarity value as the target object.

7. The method according to claim 1, characterized in that After the step of determining the target object according to the similarity value, the method includes: Obtaining the image of the target object; Inputting the image of the target object into a pre-generated visual large model to obtain the attribute data of the target object.

8. An object recognition device, characterized in that, The device includes: A data acquisition module for obtaining the voice data features of the user, the front view image of the vehicle, and the image segmentation angle; An image segmentation module for performing region segmentation on the initial front view image according to the image segmentation angle to obtain a target front view image; A feature extraction module for extracting the features of the objects in the target front view image to obtain target features; A similarity calculation module for determining a similarity value according to the target features of any object and the voice data features; An object recognition module for determining the target object according to the similarity value.

9. An electronic device, characterized in that, Including a processor and a memory, where A memory for storing a computer program; A processor for executing the program stored in the memory to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.