Methods, model training methods, devices and equipment for identifying points of interest

By acquiring initial information and relevance features of points of interest, and utilizing deep learning and intelligent question answering technologies, the accuracy and efficiency issues of determining access information for points of interest in existing technologies have been resolved, achieving rapid and accurate access information recognition.

CN114692025BActive Publication Date: 2026-04-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies cannot accurately and quickly determine the access information of points of interest, leading to navigation path errors. Furthermore, the methods of data collection and manual identification are time-consuming, labor-intensive, and inefficient.

Method used

By acquiring initial information and relevance features of the points of interest to be identified, and utilizing deep learning and intelligent question answering technologies, the access information of the points of interest can be quickly determined based on trajectory information and relevance features, reducing data collection and manual costs.

Benefits of technology

It enables the rapid and accurate determination of access information for points of interest, reducing time and labor costs and improving the accuracy and efficiency of navigation routes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692025B_ABST
    Figure CN114692025B_ABST
Patent Text Reader

Abstract

This disclosure provides methods, model training methods, devices, and equipment for identifying points of interest (POIs), relating to artificial intelligence, particularly to fields such as voice technology, map navigation, deep learning, autonomous driving, intelligent transportation, vehicle-to-everything (V2X) communication, and smart cockpits. The specific implementation involves: acquiring the trajectory information and correlation features of the POI to be identified in the current time period; the trajectory information representing the trajectory with the POI as the key location; the correlation features representing the correlation between the traffic information and trajectory information of the same POI in the same time period; the traffic information indicating the accessibility of the POI; and determining the traffic information of the POI to be identified in the current time period based on the trajectory information and correlation features. This method accurately and quickly determines the traffic information of the POI to be identified in the current time period, reducing time and labor costs and improving efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to fields such as voice technology, map navigation, deep learning, autonomous driving, intelligent transportation, vehicle networking, and intelligent cockpit in artificial intelligence, and in particular to a method for identifying points of interest, a model training method, an apparatus, and a device. Background Technology

[0002] With the development of technology, navigation technology has been provided to facilitate users' travel. When providing navigation, the accuracy of accessibility information for Points of Interest (POIs) is crucial to navigation accuracy. Accessibility information refers to information that indicates the accessibility of a POI. Accessibility information includes one or more of the following: accessible passageways to the POI, the orientation of the accessible passageways, the location of the POI, etc.

[0003] How to accurately and quickly determine the access information of points of interest is a problem that urgently needs to be solved. Summary of the Invention

[0004] This disclosure provides a method, model training method, apparatus, and device for accurately and quickly determining traffic information of points of interest.

[0005] According to a first aspect of this disclosure, a method for identifying points of interest is provided, the method comprising:

[0006] Obtain initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key location;

[0007] Obtain the relevance features of the point of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest;

[0008] Based on the trajectory information of the point of interest to be identified in the current time period and the correlation features, the traffic information of the point of interest to be identified in the current time period is determined.

[0009] According to a second aspect of this disclosure, a model training method for information recognition of points of interest is provided, the method comprising:

[0010] Obtain a set of interest points to be trained; wherein, the set of interest points to be trained includes multiple interest points to be trained, and each interest point to be trained has access information and trajectory information, wherein the access information is information that can indicate the accessibility of the interest point, and the trajectory information represents the trajectory with the interest point as the key location;

[0011] Based on the traffic information and trajectory information of the points of interest to be trained, the relevance features of the points of interest to be trained are determined; wherein, the relevance features characterize the correlation between the traffic information and trajectory information of the same point of interest in the same time period.

[0012] The relevance features of the points of interest to be trained are processed to obtain a recognition model; wherein, the recognition model is used to identify the access information of the points of interest to be identified.

[0013] According to three aspects of this disclosure, an interest point information identification device is provided, the device comprising:

[0014] The first acquisition unit is used to acquire initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key position;

[0015] The second acquisition unit is used to acquire the relevance features of the point of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest;

[0016] The determining unit is used to determine the traffic information of the point of interest to be identified in the current time period based on the trajectory information of the point of interest to be identified in the current time period and the correlation features.

[0017] According to four aspects of this disclosure, a model training apparatus for information recognition of points of interest is provided, the apparatus comprising:

[0018] The first acquisition unit is used to acquire a set of interest points to be trained; wherein, the set of interest points to be trained includes multiple interest points to be trained, and each interest point to be trained has access information and trajectory information, wherein the access information is information that can indicate the accessibility of the interest point, and the trajectory information represents the trajectory with the interest point as the key position;

[0019] The determining unit is used to determine the relevance features of the point of interest to be trained based on the traffic information and trajectory information of the point of interest to be trained; wherein, the relevance features characterize the correlation between the traffic information and trajectory information of the same point of interest in the same time period;

[0020] The training unit is used to train the relevance features of the interest points to be trained to obtain a recognition model; wherein the recognition model is used to identify the access information of the interest points to be identified.

[0021] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect or the second aspect.

[0022] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first or second aspect.

[0023] According to a seventh aspect of this disclosure, a computer program product is provided, the computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, the at least one processor executing the computer program causing the electronic device to perform the method described in the first aspect or the second aspect.

[0024] According to the eighth aspect of this disclosure, an autonomous driving vehicle is provided, wherein the autonomous driving vehicle is provided with the electronic equipment provided in the fifth aspect.

[0025] According to the technical solution of this disclosure, by acquiring the correlation features of the point of interest to be identified, the correlation features characterize the correlation between the traffic information and trajectory information of the same point of interest in the same time period. Based on the trajectory information of the point of interest to be identified in the current time period and its correlation features, the traffic information of the point of interest to be identified in the current time period can be determined. This eliminates the need for data collection vehicles or manual labor to acquire the traffic information of the point of interest to be identified in the current time period; the traffic information of the point of interest to be identified in the current time period can be accurately and quickly determined based on the correlation features, reducing time and labor costs and improving efficiency. Furthermore, for the point of interest to be identified, only the next trajectory information of the point of interest in the current time period is needed to determine its traffic information in the current time period; multiple trajectory information is not required to obtain the traffic information for the current time period, allowing for timely determination of the traffic information of the point of interest to be identified in the current time period.

[0026] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0027] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0028] Figure 1 This is a scenario that can realize the embodiments of this disclosure. Figure 1 ;

[0029] Figure 2 This is a scenario that can realize the embodiments of this disclosure. Figure 2 ;

[0030] Figure 3 This is a schematic diagram based on the first embodiment of the present disclosure;

[0031] Figure 4 This is a schematic diagram according to the second embodiment of the present disclosure;

[0032] Figure 5 This is an illustration based on the feature information provided in this disclosure. Figure 1 ;

[0033] Figure 6 This is a schematic diagram according to the third embodiment of the present disclosure;

[0034] Figure 7 This is a structural diagram of the recognition model provided in this disclosure;

[0035] Figure 8 This is a schematic diagram according to the fourth embodiment of the present disclosure;

[0036] Figure 9 This is a schematic diagram according to the fifth embodiment of the present disclosure;

[0037] Figure 10 This is an illustration based on the feature information provided in this disclosure. Figure 2 ;

[0038] Figure 11 This is a schematic diagram according to the sixth embodiment of the present disclosure;

[0039] Figure 12 This is a structural diagram of the initial model provided in this disclosure;

[0040] Figure 13 This is a schematic diagram according to the seventh embodiment of the present disclosure;

[0041] Figure 14 This is a schematic diagram according to the eighth embodiment of the present disclosure;

[0042] Figure 15 This is a schematic diagram according to the ninth embodiment of the present disclosure;

[0043] Figure 16 This is a schematic diagram according to the tenth embodiment of the present disclosure;

[0044] Figure 17 This is a schematic diagram according to the eleventh embodiment of the present disclosure;

[0045] Figure 18 A schematic block diagram of an example electronic device 1800 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0046] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0047] With the development of technology, navigation technology has been provided to facilitate users' travel. When providing navigation, the accuracy of Points of Interest (POI) information is crucial to navigation accuracy. POIs can be the destination, or they can be important locations along the user's path.

[0048] Accessibility information refers to information that indicates the accessibility of a point of interest (POI). Accessibility information includes one or more of the following: accessible passageways to the POI, the orientation of the accessible passageways to the POI, the location of the POI, etc.

[0049] For example, for a point of interest A, the access information for point of interest A is whether port a is open. Port a is either the entrance or exit of point of interest A.

[0050] If the accessibility information for a point of interest (POI) is incorrect, an accurate navigation path cannot be provided to the user or vehicle. This means the user or vehicle cannot accurately or promptly reach or pass through the POI. For example, if POI A is the destination, and its entrance 'a' is closed, but entrance 'a' is confirmed to be open, then the generated navigation path with POI A as the destination and entrance 'a' as the entrance will be incorrect, preventing the vehicle from entering POI A. Therefore, accurate and rapid determination of the accessibility information for POIs is necessary.

[0051] In one example, a data collection vehicle can be used to collect access information for points of interest (POIs). During the data collection process, many practical problems may arise. For instance, the relocation of a POI will cause its access information to be changed; if the access information for a POI becomes invalid, it will need to be collected again; if the road where the POI is located is under construction or closed, access information for that POI cannot be collected.

[0052] Therefore, the above methods require significant time and cost to collect traffic information for points of interest (POIs). Furthermore, due to numerous practical problems, timely collection of PPO traffic information is often impossible, hindering the provision of subsequent traffic information and navigation services to users or vehicles. Additionally, these methods rely on multiple trajectories to determine the current traffic information for PPOs; however, trajectories are lagging, requiring time to obtain multiple estimates, thus preventing timely and rapid acquisition of current traffic information for PPOs. In conclusion, the above methods cannot accurately and quickly determine the traffic information for PPOs.

[0053] In one example, a user's historical trajectory can be collected, which can indicate access information for points of interest. The user's historical trajectory can be manually analyzed to obtain access information for these points of interest. For instance, if a point of interest is the destination, the historical trajectory can show whether the entrance to that point is open, allowing for manual identification of access information.

[0054] However, the above methods require manual identification of numerous historical trajectories to obtain access information for each point of interest (POI), consuming significant manpower and reducing efficiency. Furthermore, manual methods are prone to errors. If a POI relocates or its access information becomes invalid, manual re-identification of historical trajectories is necessary to re-obtain access information, which takes considerable time. Additionally, these methods require manual determination of the current access information for a POI based on multiple trajectories; however, trajectories are lagging, requiring time to obtain multiple estimates, thus hindering the timely and rapid acquisition of current access information for the POI. Therefore, these methods cannot accurately and quickly determine the access information for POIs.

[0055] This disclosure provides a method, model training method, apparatus, and device for identifying points of interest, applicable to fields such as voice technology, map navigation, deep learning, autonomous driving, intelligent transportation, vehicle networking, and smart cockpits in artificial intelligence, to achieve the goal of accurately and quickly determining the access information of points of interest.

[0056] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0057] Figure 1 This is a scenario that can realize the embodiments of this disclosure. Figure 1 ,like Figure 1As shown, when a user navigates using mobile terminal 101, mobile terminal 101 needs to obtain accessibility information for points of interest (POIs). Accessibility information refers to information that indicates the accessibility of a POI. Accessibility information includes one or more of the following: accessible passageways of the POI, the orientation of the accessible passageways of the POI, the location of the POI, etc. For example, for a POI A, the accessibility information for POI A is whether passageway a is open. Passageway a is either the entrance or exit of POI A. Then, mobile terminal 101 provides a navigation path to the user based on the accessibility information of the POI, thereby providing navigation for the user.

[0058] Figure 2 This is a scenario that can realize the embodiments of this disclosure. Figure 2 ,like Figure 2 As shown, when an autonomous vehicle or user navigates using the in-vehicle terminal 201, the in-vehicle terminal 201 needs to obtain accessibility information for points of interest (POIs). Accessibility information refers to information that indicates the accessibility of a POI. Accessibility information includes one or more of the following: accessible passageways to the POI, the orientation of the accessible passageways to the POI, the location of the POI, etc. For example, for a POI A, the accessibility information for POI A is whether passageway a is open. Passageway a is either the entrance or exit of POI A. Then, the in-vehicle terminal 201 provides navigation routes to the autonomous vehicle or user based on the accessibility information of the POIs, thereby providing navigation for the autonomous vehicle or user.

[0059] Figure 3 This is a schematic diagram based on the first embodiment of the present disclosure, as shown below. Figure 3 As shown, the method for identifying points of interest provided in this embodiment includes:

[0060] S301. Obtain the initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key location.

[0061] For example, in this embodiment, the executing entity may be a server, a terminal device, an electronic device, an autonomous vehicle, a controller on an autonomous vehicle, a remote device, a point-of-interest identification device or equipment, or other devices or equipment capable of executing the method of this embodiment. This embodiment is described with an electronic device as the executing entity.

[0062] Obtain the points of interest to be identified. These points can be obtained from other devices or locally.

[0063] This embodiment requires obtaining initial information about the point of interest to be identified, including its trajectory information during the current time period. Based on this initial information, the travel information for the point of interest during the current time period can then be determined.

[0064] Accessibility information refers to information that indicates the accessibility of a point of interest (POI). Accessibility information includes one or more of the following: accessible entrances to the POI, the orientation of the accessible entrances to the POI, the location of the POI, the opening hours of the POI, etc.

[0065] For example, for a point of interest A, the accessible passage of the point of interest refers to whether the access information of point of interest A is accessed through point a. Point a is the entrance or exit of point of interest A.

[0066] For example, for a point of interest A, the orientation of the accessible passageway of the point of interest refers to the orientation of the entrance of the point of interest A; or, the orientation of the exit of the point of interest A.

[0067] For example, for a point of interest A, the location of the point of interest refers to the specific location of the entrance to the point of interest, and the specific location of the exit to the point of interest, based on the access information for point of interest A.

[0068] For example, for a point of interest A, the opening time of the point of interest refers to the access information of point of interest A, which indicates that the point of interest is open from time 1 to time 2.

[0069] Trajectory information refers to a trajectory with points of interest (POIs) as key locations. A trajectory can be a user's driving path or a vehicle's driving path. Key locations refer to POIs as either the endpoint of the trajectory or important points along the route.

[0070] For example, for a trajectory G, trajectory G ends at point of interest A. Or, for a trajectory G, an important point along the way is point of interest A.

[0071] S302. Obtain the relevance features of the points of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest.

[0072] For example, it is also necessary to obtain the relevance features of the points of interest to be identified. The relevance features indicate the correlation and association between the traffic information and trajectory information of the points of interest to be identified in the same time period.

[0073] The term "period" as used above refers to a period of time; for example, a "period" is P hours, or a "period" is P days, where P is a positive integer greater than or equal to 1.

[0074] Furthermore, the aforementioned "same time period" refers to a single time period consisting of one moment and another. If, for two time periods, the length of one time period equals the length of the other, and the two moments preceding and following each of the two time periods are the same as the two moments preceding and following each of the other time period, then these two time periods are considered the same time period. For example, 13:00 to 14:00 on January 1st constitutes time period 1; 13:00 to 14:00 on January 2nd constitutes time period 2. Time period 1 and time period 2 have the same length, and since both time period 1 and time period 2 are from 13:00 to 14:00, time period 1 and time period 2 can be considered the same time period.

[0075] Alternatively, the aforementioned "same time period" refers to a time period consisting of one moment and another. If, for two time periods, the date of one time period differs from the date of the other, then these two time periods are not the same. In this case, no two time periods are identical; each time period is distinct. For example, 13:00 to 14:00 on January 1st constitutes time period 1; 13:00 to 14:00 on January 2nd constitutes time period 2. Although time period 1 and time period 2 have the same duration—time period 1 is from 13:00 to 14:00, and time period 2 is also from 13:00 to 14:00—time period 1 is a time period following January 1st, and time period 2 is a time period following January 2nd. Time period 1 and time period 2 cannot be considered the same time period; they are two distinct and independent time periods.

[0076] S303. Based on the trajectory information of the point of interest to be identified in the current time period and the correlation characteristics, determine the passage information of the point of interest to be identified in the current time period.

[0077] For example, since the correlation feature characterizes the correlation between the traffic information and trajectory information of the same point of interest in the same time period, the traffic information of the point of interest to be identified in the current time period can be determined based on the trajectory information of the point of interest to be identified in the current time period in the initial information of the point of interest to be identified, as well as the correlation feature of the point of interest to be identified.

[0078] In this embodiment, by acquiring the relevance features of the point of interest to be identified, which characterize the correlation between the traffic information and trajectory information of the same point of interest in the same time period, the traffic information of the point of interest to be identified in the current time period can be determined based on its trajectory information and relevance features. This eliminates the need for a data collection vehicle or manual labor to acquire the traffic information of the point of interest in the current time period; the relevance features allow for accurate and rapid determination of the traffic information, reducing time and labor costs and improving efficiency. Furthermore, for the point of interest to be identified, only the next trajectory information of the point of interest in the current time period is needed to determine its traffic information, eliminating the need for multiple trajectory information sets; this allows for timely determination of the traffic information of the point of interest in the current time period.

[0079] To help readers gain a deeper understanding of the implementation principles of this disclosure, the following will be discussed in conjunction with... Figures 4-7 right Figure 3 The illustrated embodiments are further refined.

[0080] Figure 4 This is a schematic diagram based on the second embodiment of the present disclosure, as shown below. Figure 4 As shown, the method for identifying points of interest provided in this embodiment includes:

[0081] S401. Obtain the initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key location.

[0082] In one example, the initial information also includes traffic information for the point of interest to be identified in each time period within the historical time period, as well as trajectory information for the point of interest to be identified in each time period within the historical time period; the point of interest to be identified has relevance characteristics in each time period within the historical time period. The historical time period includes at least one time period.

[0083] In one example, access information includes one or more of the following: accessible passageway, the orientation of the accessible passageway, the location of the point of interest, and the opening hours of the point of interest.

[0084] For example, in this embodiment, the executing entity may be a server, a terminal device, an electronic device, an autonomous vehicle, a controller on an autonomous vehicle, a remote device, a point-of-interest identification device or equipment, or other devices or equipment capable of executing the method of this embodiment. This embodiment is described with an electronic device as the executing entity.

[0085] Initial information about the point of interest to be identified can be obtained from other devices or databases. This initial information includes the trajectory information of the point of interest to be identified during the current time period.

[0086] Trajectory information refers to a trajectory with points of interest (POIs) as key locations. A trajectory can be a user's driving path or a vehicle's driving path. Key locations refer to POIs as either the endpoint of the trajectory or important points along the route.

[0087] To more accurately identify the traffic information of the point of interest to be identified in the current time period, it is also necessary to obtain the traffic information and trajectory information of the point of interest to be identified in each time period within a historical time period. The "current time period" mentioned above is not included in the historical time period.

[0088] It can be seen that, for the point of interest to be identified, the traffic information of the point of interest in each time period within the historical time period is the actual traffic information of the point of interest in each time period within the historical time period. Furthermore, for the point of interest to be identified, the trajectory information of the point of interest in each time period within the historical time period is the actual trajectory information of the point of interest in each time period within the historical time period.

[0089] The access information for a point of interest (POI) within a historical time period includes both positive and negative samples. Different weights can be assigned to these positive and negative samples to ensure a balanced information profile for the same POI across different time periods. For example, access information might indicate whether the entry point to the POI is open. Therefore, for a given POI, access information in one historical time period represents an open entry point (positive sample), while access information in the next historical time period represents an closed entry point (negative sample).

[0090] For example, the historical time period includes three time periods: Time Period 1, Time Period 2, and Time Period 3. The access information for the Point of Interest (POI1) to be identified in Time Period 1 indicates that the entrance to POI1 is open during Time Period 1; and POI1 has trajectory information within Time Period 1. The access information for the Point of Interest (POI1) to be identified in Time Period 2 indicates that the entrance to POI1 is not open during Time Period 2; and POI1 has trajectory information within Time Period 2. The access information for the Point of Interest (POI1) to be identified in Time Period 3 indicates that the entrance to POI1 is open during Time Period 3; and POI1 has trajectory information within Time Period 3.

[0091] Access information includes one or more of the following: accessible passageways, the orientation of accessible passageways, the location of the point of interest (POI), and the opening hours of the POI. Access information can include multiple types of information; thus, multiple access information for the POI to be identified can be determined.

[0092] In the above process, the traffic information of the point of interest to be identified in each time period within the historical time period is real traffic information; this information is determined based on the historical trajectory with the point of interest to be identified as the key point. Alternatively, the traffic information of the point of interest to be identified in each time period within the historical time period can be obtained based on intelligent question answering; see the introduction below.

[0093] In one example, the access information for the point of interest to be identified in each time period within the historical time period is determined based on each user's voice; wherein, the content of the user's voice includes the access information for the point of interest to be identified in each time period within the historical time period; the user's voice is obtained based on intelligent question answering.

[0094] In one example, the traffic information of the point of interest to be identified in each time period within the historical time period is the traffic information in each corrected speech; wherein, the corrected speech is the candidate speech with the highest value represented by the similarity information; the similarity information is the similarity between the user's speech and the candidate speech in the preset candidate speech set; the preset candidate speech set includes at least one candidate speech.

[0095] In one example, the similarity information is the similarity between the first audio information and the second audio information; wherein, the first audio information is the audio information of the user's speech, and the second audio information is the audio information of the candidate speech.

[0096] For example, in this embodiment, the access information of the point of interest to be identified in each time period within a historical time period can be obtained based on intelligent question answering.

[0097] If the actual traffic information of the point of interest to be identified for each time period in the historical time period has not been obtained before, or if the location of the point of interest to be identified is relatively marginal and it is difficult to determine the traffic information of the point of interest to be identified for each time period in the historical time period through multiple trajectories, the traffic information of the point of interest to be identified for each time period in the historical time period can be obtained through intelligent question answering based on voice interaction.

[0098] When obtaining traffic information for a point of interest (POI) within a historical time period using intelligent question answering, multiple texts are pre-set, each containing a question about the POI. A Text-to-Speech (TTS) model is used to extract text from the pre-set texts; then, based on the TTS model, the extracted text is used to generate the spoken question.

[0099] The generated question message is then sent to the user. For example, a phone call is made to the user, and then the question message is sent.

[0100] This allows users to reply with intelligent voice questions, and in turn, the system can obtain the user's voice messages, which contain traffic information for the time period in which the point of interest to be identified is located.

[0101] For example, based on a TTS (Text-to-Speech) model, text is converted into human-like spoken questions. A virtual digital human then sends these spoken questions to the user via telephone. Examples of spoken questions include: "Is there parking available in front of the store?", "Are there any parking lots nearby?", "Does the scenic area have traffic restrictions during holidays?", "Does your store face Road A?", and so on. The user can then respond with spoken responses. This allows the acquisition of the user's spoken responses, such as: "Parking is available in front of the store", "There are no parking lots nearby", "The scenic area has traffic restrictions during holidays", "The store faces Road A", and so on.

[0102] By acquiring user voice recordings containing traffic information through intelligent question-and-answer methods, and then extracting traffic information for specific time periods within a historical timeframe from the user voice recordings, the system can quickly obtain traffic information for every time period within a historical timeframe for the target point of interest. This eliminates the need for data collection vehicles or manual methods to acquire traffic information for specific time periods within a historical timeframe. It also facilitates the rapid subsequent identification of the traffic information for the target point of interest in the current time period.

[0103] In one example, multiple preset question texts are pre-set, each containing text information about a question asked regarding a point of interest to be identified. Based on a TTS model, the preset question texts are extracted from the preset texts; then, based on the extracted preset question texts, the spoken question is generated.

[0104] In intelligent question-and-answer interactions with users, to achieve smooth dialogue and effectively reduce the probability and rate of users hanging up, user intent can be obtained. Users can issue a voice message upon first speaking, or after receiving the first voice question. Then, based on the user's voice message, intent recognition methods are used to identify the user's intent.

[0105] Then, a new question voice is generated based on the user's intent. At this point, it is necessary to determine a preset question text that corresponds to both the current point of interest to be identified and the user's intent, based on the traffic information to be obtained for the current point of interest to be identified. For example, based on the correspondence between the point of interest to be identified, the user's intent, and the preset question text, a preset question text is determined that corresponds to both the current point of interest to be identified and the user's intent. Thus, a new question voice is generated based on the determined preset question text. Alternatively, it is necessary to determine a preset question text that corresponds to the current point of interest to be identified, based on the traffic information to be obtained for the current point of interest to be identified. For example, based on the correspondence between the point of interest to be identified, the preset question text is determined that corresponds to the current point of interest to be identified. Furthermore, based on the current user intent and the preset question text corresponding to the current point of interest to be identified, a complete and fluent question voice is generated, thus obtaining a new question voice.

[0106] A new voice message with a question is sent via telephone. Since the user will provide feedback based on this new voice message, the user's voice messages can be obtained.

[0107] Through the aforementioned intelligent voice interaction process, questions can be uttered to the user based on their intent, inquiring about the accessibility information of the point of interest to be identified during specific time periods within a historical timeframe. This enables smooth dialogue with the user and effectively reduces the probability and rate of users hanging up; thus, it allows for the effective, accurate, and rapid determination of the accessibility information of the point of interest to be identified during each time period within a historical timeframe.

[0108] For example, based on a preset question text, a question voice is generated: "Is this location A?"; location A is a point of interest to be identified. Using an intelligent digital human, the generated question voice is sent to the user via telephone. The user's voice is received: "Yes, who is this?". At this point, it's necessary to determine the user's intent to generate a suitable new question voice to reduce the probability and proportion of the user hanging up. Based on the user's voice "Yes, who is this?", the user's intent can be extracted as "The user has questions about the current call content." Then, based on the user's intent and the preset question text, a new question voice is generated: "I am intelligent customer service B. Is entrance A of the store currently open?" Then, using the intelligent digital human, the generated new question voice is sent to the user via telephone. When the user responds with the new voice "Entrance A is open," the access information for the point of interest to be identified can be extracted: "Entrance A is currently open." Note that the access information extracted here is the access information for the point of interest during a specific time period in history.

[0109] Furthermore, when extracting access information corresponding to a user's voice—that is, when extracting access information from a user's voice—the extracted access information may be inaccurate due to the influence of the user's accent. The extracted access information may not reflect the true meaning conveyed by the user's voice. Therefore, it is necessary to correct the user's voice based on their accent.

[0110] A preset candidate speech set is provided, which includes at least one candidate speech, which can be in Mandarin pronunciation; the content of the candidate speech is traffic information. This preset candidate speech set needs to be extracted first.

[0111] Alternatively, multiple preset candidate speech sets can be established, each containing at least one candidate speech, the content of which is travel information. Each candidate speech set corresponds to a geographical region. For each candidate speech set, the candidate speech within that set is spoken in the geographical region corresponding to that set. Thus, for the current user's speech, the geographical region of the user is known in advance; consequently, the preset candidate speech set matching the user's speech can be extracted.

[0112] For example, candidate voice information can include: location information, orientation information, whether parking is available, road name, whether the entrance is open, whether the exit is open, the name of the point of interest, the names of adjacent points of interest, and so on.

[0113] Then, the similarity between the user's voice and each candidate voice in the extracted preset candidate voice set is calculated to obtain the similarity information of each candidate voice.

[0114] In one example, during the process of obtaining similarity information for each candidate speech, if speech recognition and text processing are performed on the user's speech to obtain the corresponding text, and then the similarity between the text corresponding to the user's speech and the text corresponding to each candidate speech is calculated, this method is inaccurate. This is because if the user's speech has a regional accent, the text obtained from speech recognition and text processing will be inaccurate. Therefore, this method cannot be used to determine similarity information.

[0115] In this embodiment, during the process of obtaining the similarity information of each candidate speech, the similarity between the user's speech and each candidate speech in the extracted preset candidate speech set can be determined based on the audio.

[0116] One approach involves: performing audio analysis on the user's speech to obtain the first audio information; then performing audio analysis on each candidate speech to obtain the second audio information. For each candidate speech, the similarity between the first audio information of the user's speech and the second audio information of the candidate speech is calculated, yielding the similarity information cos=(X1*X2) / (norm(X1)*norm(X2)). Here, X1 is the first audio information of the user's speech, X2 is the second audio information of the candidate speech, and norm() is a function that assigns length and size to vectors in a vector space.

[0117] Another approach involves performing audio analysis on the user's speech to obtain the first audio information of the user's speech; then performing audio analysis on each candidate speech to obtain the second audio information of each candidate speech. For each candidate speech, the first audio information of the user's speech and the second audio information of the candidate speech are input into a preset binary classification model; then, based on the preset binary classification model, the audio similarity between the first audio information of the user's speech and the second audio information of the candidate speech is calculated to obtain the similarity information of the candidate speech.

[0118] By calculating the audio similarity between the first audio information of the user's speech and the second audio information of each candidate speech, the similarity information of each candidate speech can be accurately determined. This is beneficial for subsequently identifying the candidate speech most similar to the user's speech, and thus accurately correcting the user's speech.

[0119] After obtaining the similarity between the user's voice and each candidate voice in the extracted preset candidate voice set, the candidate voice with the highest similarity value is determined as the corrected voice. It can be seen that since the content of the candidate voice has common information, and the corrected voice is a candidate voice in the extracted preset candidate voice set, the corrected voice has common information.

[0120] Then, the passage information in the corrected speech can be identified as the passage information of the user's speech. This allows for the accurate extraction of the passage information from the user's speech.

[0121] For the point of interest to be identified, the user voice corresponding to the point of interest to be identified obtained based on the intelligent question answering method can be used to determine the passage information in the user voice in step S401, and then the passage information of the point of interest to be identified in each time period in the historical period can be extracted.

[0122] Furthermore, for each point of interest to be identified in each time period, the user's voice for that point of interest in each time period is obtained based on an intelligent question-and-answer method. Then, based on step S401, the traffic information of the point of interest to be identified in each time period within the historical time period is determined.

[0123] In the above process, the trajectory information of the point of interest to be identified for each time period within the historical time period is real traffic information. This trajectory information can be obtained from other devices or databases.

[0124] S402. Obtain the relevance features of the points of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; access information is information that can indicate the accessibility of the point of interest.

[0125] In one example, the point of interest to be identified has relevance features for each time period within a historical time period; wherein the historical time period includes at least one time period.

[0126] For example, the point of interest to be identified has access information and trajectory information for each time period within a historical time period. Then, for each time period within the historical time period, the access information and trajectory information of the point of interest to be identified are analyzed and processed to determine the correlation information between the access information and trajectory information of the point of interest to be identified within the same time period, thereby obtaining the correlation characteristics of the point of interest to be identified for each time period within the historical time period. These correlation characteristics indicate the correlation and association relationships between the access information and trajectory information of the point of interest to be identified within the same time period.

[0127] The term "time period" as used above refers to a period of time; for example, a "time period" is P hours, or a "time period" is P days, where P is a positive integer greater than or equal to 1. For an explanation of "time period," please refer to step S302, which will not be repeated here.

[0128] For example, the historical time period includes time period 1 and time period 2. For a point of interest (POI) A to be identified, this embodiment can be used to obtain the traffic information and trajectory information of POI A in time period 1. Then, based on the traffic information and trajectory information of POI A in time period 1, a correlation feature 1 for POI A in time period 1 can be determined. Correlation feature 1 characterizes the correlation and association between the traffic information and trajectory information of POI A in time period 1. Similarly, this embodiment can be used to obtain the traffic information and trajectory information of POI A in time period 2. Then, based on the traffic information and trajectory information of POI A in time period 2, a correlation feature 2 for POI A in time period 2 can be determined. Correlation feature 2 characterizes the correlation and association between the traffic information and trajectory information of POI A in time period 2. This process continues, resulting in the correlation features of the POI for each time period within the historical time period. Note that the traffic information, trajectory information, and correlation features in this example are historical information of the points of interest to be identified, not information in the current time period.

[0129] S403. Obtain the feature image of the interest point to be identified; wherein, the feature image is used to indicate the features of the interest point.

[0130] For example, the traffic information of the point of interest to be identified in the current time period can be obtained based on the feature image of the point of interest to be identified and the correlation features of the point of interest. The feature image of the point of interest to be identified can indicate the characteristics of the point of interest. Thus, more features can be added to obtain the traffic information of the point of interest to be identified in the current time period.

[0131] First, a feature image of the point of interest to be identified needs to be obtained. The feature image is generated based on the point of interest information. This information includes one or more of the following: the specific location of the point of interest, the road network conditions corresponding to the point of interest, the driving trajectory corresponding to the point of interest, etc. Thus, the feature image can indicate the characteristics of the point of interest to be identified.

[0132] Then, the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the initial information of the point of interest to be identified can be processed to obtain the traffic information of the point of interest to be identified in the current time period.

[0133] Therefore, by incorporating more features, we can obtain the traffic information of the point of interest to be identified in the current time period. This allows for more accurate identification of the traffic information of the point of interest in the current time period.

[0134] In one example, step S403 includes the following steps:

[0135] The first step of step S403 is to obtain the interest point information of the interest point to be identified; wherein the interest point information includes at least one feature information.

[0136] In the second step of step S403, if at least one feature information includes road network information and trajectory data information; wherein the road network information is the road network information of the location of the point of interest; and the trajectory data information is the trajectory data information with the point of interest as the key location, then the trajectory data information is subjected to deviation repair processing based on the road network information to obtain the repaired trajectory data information.

[0137] The third step of step S403 is to perform data fusion based on the feature information in the interest point information of the interest point to be identified from the data source, and obtain the feature image of the interest point to be identified.

[0138] In one example, at least one feature includes one or more of the following: location information, map information, road network information, trajectory data information, and areal area information.

[0139] Among them, location information is the location information of the point of interest; map information is the map information of the location of the point of interest; road network information is the road network information of the location of the point of interest; trajectory data information is the trajectory data information with the point of interest as the key location; and areal region information is the areal region information of the point of interest.

[0140] For example, the following process can be used to implement step S403.

[0141] First, it is necessary to obtain the feature image of the point of interest to be identified. The feature image is generated based on the point of interest information of the point of interest to be identified; the feature image can indicate the features of the point of interest to be identified. Therefore, it is necessary to obtain the point of interest information of the point of interest to be identified so that the feature image of the point of interest to be identified can be generated based on the point of interest information.

[0142] The interest point information to be identified includes at least one feature. The feature information can be any one of location information, map information, road network information, trajectory data information, or areal region information.

[0143] Location information refers to the location information of a point of interest; location information can be the specific location of the area, road, or intersection where the point of interest is located, or it can be the latitude and longitude of the point of interest.

[0144] Map information refers to the map information of the location of the point of interest; map information includes detailed map content matching the point of interest.

[0145] Road network information refers to the road network information of the location of a point of interest; road network information can be obtained from the road network system.

[0146] Trajectory data refers to trajectory data with points of interest (POIs) as key locations. A key location refers to a POI being either the endpoint of the trajectory or an important point along the trajectory. Trajectory data can be a user's driving trajectory or a vehicle's driving trajectory.

[0147] Area information refers to the area information of points of interest. Area information includes one or more of the following: Area of ​​Interest (AOI) area information and Building Unknown (BUD) area information. The "Area of ​​Interest" can also be called the "Area of ​​Interest".

[0148] The above process can obtain a variety of feature information, making the interest point to be identified feature complete; then, by integrating at least one feature information from the interest point information of the interest point to be identified into the relevance feature of the interest point to be identified and the initial information of the interest point to be identified, the traffic information of the interest point to be identified in the current time period can be accurately obtained.

[0149] For example, Figure 5 This is an illustration based on the feature information provided in this disclosure. Figure 1 ,like Figure 5 As shown, for a Point of Interest (POI) to be identified, the following feature maps can be obtained: BUD (Browser Undefined Area) region information 501, AOI (Area Undefined Area) region information 502, current interest point information 503, road network information 504, trajectory link 505, and trajectory data information 506. The trajectory link refers to the corrected trajectory data obtained by performing deviation correction processing on the trajectory data based on road network information. Current interest point information refers to the location information, name, etc., of the interest point. The pixel parameters of the feature map for each feature information can be 60*60; the physical world size covered by the feature map of each feature information is 80 meters * 80 meters. For the point of interest (POI1) to be identified, the above-mentioned feature information of the point of interest (POI1) to be identified is fused using the method of this embodiment to obtain the feature image 507 of the point of interest (POI1) to be identified.

[0150] Then, after obtaining the point of interest (POI) information, if the POI information includes road network information and trajectory data, the trajectory data can be repaired. Road network information can characterize the true and accurate road conditions at the location of the POI. However, trajectory data is obtained from the user's or vehicle's driving process, and deviations may occur during this process, leading to potentially inaccurate trajectory data, or the inclusion of redundant information.

[0151] Therefore, deviation correction processing can be performed on trajectory data based on road network information. In one example, a Hidden Markov Model (HMM) can be used to process the trajectory data based on road network information. This corrects the inherent deviation between the road network information and the trajectory data, making the trajectory data closer to the road network information, and then outputting the corrected trajectory data. This deviation correction processing based on road network information avoids the introduction of noise into the feature image due to trajectory data offset.

[0152] Next, the feature information in the interest point information of the points of interest to be identified needs to be fused to obtain the feature image of the points of interest to be identified. A data source-based approach can be used to fuse the feature information in the interest point information of the points of interest to be identified, rather than using the RGB (Red, Green, Blue) method of the image to fuse the feature information.

[0153] Data fusion of feature information from the interest point information to be identified based on the data source allows for relatively independent fusion of feature information. One type of feature information will not interfere with another, resulting in lossless data representation of the information in the feature image of the interest point to be identified. That is, it overcomes the data loss problem that occurs when fusing feature information based on RGB. The information in the feature image of the interest point to be identified is accurate and uninterrupted.

[0154] Based on the feature information in the interest point information of the interest point to be identified, a feature image of the interest point to be identified is generated; then the feature image of the interest point to be identified can indicate the features of the interest point to be identified.

[0155] S404. Process the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period to obtain the traffic information of the point of interest to be identified in the current time period.

[0156] In one example, step S404 includes: processing the feature image of the point of interest to be identified, the relevance features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period based on the recognition model to obtain the traffic information of the point of interest to be identified in the current time period.

[0157] For example, the above process obtains the correlation features of the point of interest to be identified in each time period within the historical time period and the trajectory information of the point of interest to be identified in the current time period. Then, based on the correlation features of the point of interest to be identified in each time period within the historical time period, the trajectory information of the point of interest to be identified in the current time period can be processed to obtain the traffic information of the point of interest to be identified in the current time period.

[0158] Furthermore, during the processing, feature images of the points of interest to be identified can be added; thus, based on the correlation features of the points of interest to be identified in each time period within the historical time period and the feature images of the points of interest to be identified, the trajectory information of the points of interest to be identified in the current time period is processed to obtain the traffic information of the points of interest to be identified in the current time period.

[0159] In one example, based on a trained recognition model, the trajectory information of the point of interest to be identified in the current time period can be processed according to the correlation features of the point of interest to be identified in each time period in the historical time period and the feature image of the point of interest to be identified, so as to obtain the traffic information of the point of interest to be identified in the current time period.

[0160] Based on the recognition model, the traffic information of the point of interest to be identified in the current time period can be obtained automatically and accurately. The recognition model can be obtained using the "Model Training Method for Information Recognition of Points of Interest" provided in this disclosure.

[0161] In one example, step S404 can be implemented using the following steps:

[0162] The first step of step S404 is to input the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period into the first fully connected layer in the recognition model for feature fusion to obtain the fused features of the point of interest to be identified.

[0163] The second step of step S404 is to input the fused features of the interest point to be identified into the second fully connected layer in the identification model for prediction processing to obtain the traffic information of the interest point to be identified in the current time period.

[0164] For example, step S404 can be implemented in the following manner.

[0165] The feature image of the point of interest to be identified, its correlation features, and its trajectory information in the current time period can be input into the recognition model. Based on the recognition model, the traffic information of the point of interest to be identified in the current time period can be predicted. The recognition model includes two fully connected layers (FC), namely the first fully connected layer and the second fully connected layer.

[0166] The feature image of the point of interest to be identified, the correlation features of the point of interest to be identified in each time period of the historical time period, and the trajectory information of the point of interest to be identified in the current time period are input into the first fully connected layer in the recognition model. Based on the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified in each time period of the historical time period, and the trajectory information of the point of interest to be identified in the current time period, feature fusion is performed to obtain the fused features of the point of interest to be identified.

[0167] For example, feature fusion can be performed on the feature image of the point of interest to be identified and its trajectory information in the current time period based on the first fully connected layer to obtain intermediate information, thereby enriching the information in the trajectory information of the point of interest to be identified in the current time period. Then, feature fusion and comparative analysis are performed on the intermediate information and the correlation features of the point of interest to be identified in each time period in the historical time period to determine the correlation features related to the intermediate information. Thus, the fused features of the point of interest to be identified are obtained, which are the correlation features related to the intermediate information.

[0168] Alternatively, for example, the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified in each time period of the historical time period, and the trajectory information of the point of interest to be identified in the current time period can be fused based on the first fully connected layer to obtain the fused features of the point of interest to be identified.

[0169] Then, the fused features of the interest points to be identified are input into the second fully connected layer in the identification model; based on the fused features of the interest points to be identified in the second fully connected layer, the traffic information is predicted to obtain the traffic information of the interest points to be identified in the current time period.

[0170] If the fusion features of the point of interest to be identified, based on the output of the first fully connected layer, are correlation features related to intermediate information, then the fusion features of the point of interest to be identified can indicate the correlation and relevance between trajectory information and traffic information. Therefore, after processing the fusion features of the point of interest to be identified and the trajectory information of the point of interest to be identified in the current time period based on the second fully connected layer, the traffic information of the point of interest to be identified in the current time period can be output.

[0171] If the fusion features of the point of interest to be identified, based on the output of the first fully connected layer, are obtained by fusing the feature image of the point of interest, the correlation features of the point of interest in each time period in the historical time period, and the trajectory information of the point of interest in the current time period, then the fusion features of the point of interest can indicate the correlation and relevance between the trajectory information and the traffic information. Furthermore, the fusion features of the point of interest include the trajectory information of the point of interest in the current time period. Then, the fusion features of the point of interest can be further processed based on the second fully connected layer to directly output the traffic information of the point of interest in the current time period.

[0172] This system utilizes multiple fully connected layers to fuse and predict the feature image of the point of interest to be identified, the relevance features of the point of interest in each historical time period, and the trajectory information of the point of interest in the current time period. This allows for better integration of the feature image of the point of interest to be trained and the trajectory information of the point of interest in the current time period, and also provides a better understanding of the relevance features of the point of interest in the current time period. The use of multiple fully connected layers enables non-linear processing of these features, including the feature image of the point of interest to be identified, the relevance features of the point of interest in each historical time period, and the trajectory information of the point of interest in the current time period.

[0173] In this embodiment, based on the above embodiments, user voice is acquired using an intelligent question-and-answer method. Traffic information for each time period within a historical timeframe is extracted from the user voice, along with trajectory information for each time period within the historical timeframe. This process yields each point of interest to be identified. The intelligent question-and-answer method allows for rapid acquisition of traffic information for each time period within a historical timeframe, reducing time and labor costs and accelerating the recognition process. Furthermore, the acquired user voice can be corrected to obtain accurate traffic information for each time period within the historical timeframe. Based on the traffic and trajectory information for each time period within the historical timeframe, the relevance features of the point of interest are determined. These relevance features characterize the correlation between traffic and trajectory information for the same point of interest within the same time period. Based on the relevance features and feature images of the point of interest within the historical timeframe, traffic information for the current time period is predicted. This incorporates more features to predict the traffic information for the point of interest in the current time period, improving recognition accuracy and precision. Furthermore, based on the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period, the traffic information of the point of interest to be identified in the current time period can be accurately and automatically determined.

[0174] Figure 6 This is a schematic diagram based on the third embodiment of the present disclosure, as shown below. Figure 6 As shown, the method for identifying points of interest provided in this embodiment includes:

[0175] S601. Obtain the initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key location.

[0176] In one example, the initial information also includes traffic information for the point of interest to be identified in each time period within the historical time period, as well as trajectory information for the point of interest to be identified in each time period within the historical time period; the point of interest to be identified has relevance characteristics in each time period within the historical time period. The historical time period includes at least one time period.

[0177] In one example, access information includes one or more of the following: accessible passageway, the orientation of the accessible passageway, the location of the point of interest, and the opening hours of the point of interest.

[0178] For example, in this embodiment, the executing entity may be a server, a terminal device, an electronic device, an autonomous vehicle, a controller on an autonomous vehicle, a remote device, a point-of-interest identification device or equipment, or other devices or equipment capable of executing the method of this embodiment. This embodiment is described with an electronic device as the executing entity.

[0179] This step is the same as step S401 and will not be repeated here.

[0180] S602. Obtain the relevance features of the points of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; access information is information that can indicate the accessibility of the point of interest.

[0181] In one example, the point of interest to be identified has relevance features for each time period within a historical time period; wherein the historical time period includes at least one time period.

[0182] For example, this step is the same as step S402, and will not be described again.

[0183] S603. Obtain the feature image of the interest point to be identified; wherein, the feature image is used to indicate the features of the interest point.

[0184] For example, this step is the same as step S403, and will not be described again.

[0185] S604. Input the feature image of the interest point to be identified into at least one convolutional neural network unit in the recognition model for processing to obtain the processed feature image of the interest point to be identified.

[0186] In one example, at least one convolutional neural network unit includes Q first convolutional neural network sub-units and MQ second convolutional neural network sub-units. The structure of the first convolutional neural network sub-units differs from the structure of the second convolutional neural network sub-units. Q is a positive integer greater than 1, M is a positive integer greater than 1, and M is greater than Q. Step S604 includes the following steps:

[0187] The first step of step S604. Repeat the following steps until the first preset condition is met: input the feature image of the interest point to be identified into the qth first convolutional neural network subunit in the recognition model for processing to obtain the first feature map of the interest point to be identified; determine the first feature map of the interest point to be identified as the new feature image of the interest point to be identified, and determine that the value of q is q+1; where the initial value of q is 1, and q is a positive integer greater than or equal to 1 and less than or equal to Q; the new feature image obtained when the first preset condition is met is the intermediate feature of the interest point to be identified.

[0188] The second step of step S604. Repeat the following steps until the second preset condition is met: input the intermediate features of the interest point to be identified into the kth second convolutional neural network subunit in the recognition model for processing to obtain the second feature map of the interest point to be identified; determine that the second feature map of the interest point to be identified is the new intermediate feature, and determine that the value of k is k+1; where the initial value of k is 1, and k is a positive integer greater than or equal to 1 and less than or equal to MQ.

[0189] Among them, the intermediate features obtained when the second preset condition is met are the processed feature images of the interest points to be identified.

[0190] For example, after obtaining the feature image of the point of interest to be identified, the feature image needs to be processed based on the recognition model.

[0191] Figure 7 Based on the structural diagram of the recognition model provided in this disclosure, as shown below. Figure 7 As shown, the recognition model 700 includes an input layer 701, at least one convolutional neural network unit 702, a long short-term memory (LSTM) network unit 703, a first fully connected layer 704, a second fully connected layer 705, and an output layer 706.

[0192] For a point of interest to be identified, its feature image is input to at least one convolutional neural network unit 702 through the input layer 701. The feature image of the point of interest is processed by the at least one convolutional neural network unit 702 to obtain a processed feature image of the point of interest, thereby completing the processes of feature extraction and feature compression. The resulting processed feature image of the point of interest can be better combined with the relevant features of the point of interest.

[0193] For the point of interest to be identified, the LSTM unit 703 processes the traffic information and trajectory information of the point of interest to be identified in each time period to obtain the correlation features of the point of interest to be identified in each time period.

[0194] For example, the historical time period includes three time periods: Time Period 1, Time Period 2, and Time Period 3. For a Point of Interest (POI) a to be identified, based on the traffic and trajectory information of POI a in Time Period 1 of the historical time period, the relevance features of POI a in Time Period 1 of the historical time period are obtained; based on the traffic and trajectory information of POI a in Time Period 2 of the historical time period, the relevance features of POI a in Time Period 2 of the historical time period are obtained. Then, based on the LSTM unit 703, the relevance features of POI a in each time period of the historical time period are obtained.

[0195] For the point of interest to be identified, the correlation features of the point of interest in each time period in the historical time period, the processed feature image of the point of interest, and the trajectory information of the point of interest in the current time period are input into the first fully connected layer 704 and the second fully connected layer 705 for processing in sequence; and the passage information of the point of interest to be identified in the current time period can be output through the output layer 706.

[0196] In one example, the technical solution for step S604 can be implemented in the following way.

[0197] The recognition model includes at least one convolutional neural network (CNN) unit; the number of CNN units is M, where M is a positive integer greater than 1. Furthermore, the M CNN units include Q first CNN sub-units and MQ second CNN sub-units, where Q is a positive integer greater than 1 and less than M. The structure of the first CNN sub-units differs from that of the second CNN sub-units, and the feature images of the points of interest to be recognized are processed based on these different CNN unit structures.

[0198] After acquiring the feature image of the point of interest to be identified, the feature image is input into the first convolutional neural network subunit. The first convolutional neural network subunit processes the feature image of the point of interest to be identified, obtaining a first feature map of the point of interest. Then, this first feature map is used as the new feature image of the point of interest. The new feature image of the point of interest is then input into the second convolutional neural network subunit. The second convolutional neural network subunit processes the new feature image of the point of interest to be identified, obtaining a new first feature map of the point of interest. Then, this first feature map is used as the new feature image of the point of interest. This process continues until the new feature image of the point of interest is input into the q-th convolutional neural network subunit. The q-th convolutional neural network subunit processes the new feature image of the point of interest to be identified, obtaining a new first feature map of the point of interest. Then, this first feature map is used as the new feature image of the point of interest, and the value of q is incremented by 1. q is a positive integer greater than or equal to 1 and less than or equal to Q. This process continues until the first preset condition is met.

[0199] The first preset condition can be "determining that the value of q is equal to Q". Alternatively, the first preset condition can be receiving a first stop instruction, which is used to instruct the generation of the first feature map of the interest point to be identified to stop.

[0200] The feature image obtained when the first preset condition is met is used as the intermediate feature of the interest point to be identified.

[0201] Then, after obtaining the intermediate features of the interest point to be identified, these intermediate features are input into the first second convolutional neural network subunit. Based on the feature image of the interest point to be identified by the first second convolutional neural network subunit, a second feature map of the interest point to be identified is obtained. This second feature map is then used as a new intermediate feature of the interest point to be identified. The new intermediate features of the interest point to be identified are then input into the second second convolutional neural network subunit. Based on the new feature image of the interest point to be identified by the second second convolutional neural network subunit, a new second feature map of the interest point to be identified is obtained. This second feature map is then used as a new intermediate feature of the interest point to be identified. This process continues, with the new intermediate features of the interest point to be identified being input into the kth second convolutional neural network subunit. Based on the new feature image of the interest point to be identified by the kth second convolutional neural network subunit, a new second feature map of the interest point to be identified is obtained. This second feature map is then used as a new intermediate feature of the interest point to be identified, and the value of k is incremented by 1. k is a positive integer greater than or equal to 1 and less than or equal to MQ. This process continues until the second preset condition is met.

[0202] The second preset condition can be "determining that the value of k is equal to MQ". Alternatively, the second preset condition can be receiving a second stop instruction, which is used to indicate the cessation of generating the second feature map of the interest point to be identified.

[0203] The intermediate features obtained when the second preset condition is met are used as the processed feature image of the interest points to be identified.

[0204] Therefore, the feature images of the points of interest to be identified are processed based on convolutional neural network units with different structures. For example, the first convolutional neural network subunit has fewer structural layers than the second convolutional neural network subunit; or, the second convolutional neural network subunit has fewer structural layers than the first convolutional neural network subunit; thereby reducing the computational complexity of the recognition process and speeding up the recognition process.

[0205] In one example, the q-th first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer. The first step of step S604, "inputting the feature image of the interest point to be identified into the q-th first convolutional neural network subunit in the recognition model for processing to obtain the first feature map of the interest point to be identified," includes the following process: inputting the feature image of the interest point to be identified into the first convolutional layer for convolution processing to obtain the first convolutional feature; wherein, the first convolutional feature represents the convolutional feature of the feature image of the interest point to be identified; inputting the first convolutional feature into the first activation layer for correction processing to obtain the corrected first convolutional feature; inputting the corrected first convolutional feature into the pooling layer for feature compression processing to obtain the first feature map of the interest point to be identified; wherein, the first feature map of the interest point to be identified represents the feature image of the interest point to be identified.

[0206] The k-th second convolutional neural network subunit includes a second convolutional layer and a second activation layer. The second step of step S604, "inputting the intermediate features of the interest point to be identified into the k-th second convolutional neural network subunit in the recognition model for processing to obtain the second feature map of the interest point to be identified," includes the following process: inputting the intermediate features of the interest point to be identified into the second convolutional layer for convolution processing to obtain the second convolutional features; wherein, the second convolutional features represent the convolutional features of the intermediate features of the interest point to be identified; inputting the second convolutional features into the second activation layer for correction processing to obtain the second feature map of the interest point to be identified.

[0207] For example, the technical solution of step S604 can be implemented in the following manner.

[0208] The recognition model includes at least one convolutional neural network unit; the number of at least one convolutional neural network unit is M, where M is a positive integer greater than 1. Furthermore, the M convolutional neural network units include Q first convolutional neural network sub-units and MQ second convolutional neural network sub-units, where Q is a positive integer greater than 1 and Q is less than M.

[0209] Each first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer. However, each second convolutional neural network subunit only includes a second convolutional layer and a second activation layer. Therefore, the second convolutional neural network subunit does not require a pooling layer.

[0210] After acquiring the feature image of the interest point to be identified, it is input into the first convolutional neural network subunit. Based on the first convolutional layer of the first convolutional neural network subunit, convolution calculations are performed on the feature image of the interest point to be identified, resulting in the first convolutional feature. This first convolutional feature is the convolutional feature of the feature image of the interest point to be identified. Then, based on the first activation layer of the first convolutional neural network subunit, the first convolutional feature is modified to remove redundant features; normalization may also be performed, resulting in a modified first convolutional feature. Next, based on the pooling layer of the first convolutional neural network subunit, the modified first convolutional feature is compressed, reducing its dimensionality. For example, if the modified first convolutional feature has a dimension of 100*100, the pooling layer reduces its dimension to 50*50. This yields the first feature map of the interest point to be identified. This first feature map is then used as the new feature image of the interest point to be identified.

[0211] The new feature image of the interest point to be identified is input into the second first convolutional neural network subunit. Based on the first convolutional layer of the second first convolutional neural network subunit, convolution calculation is performed on the new feature image of the interest point to be identified to obtain the first convolutional feature. It can be seen that the first convolutional feature is the new convolutional feature of the feature image of the interest point to be identified. Then, based on the first activation layer of the second first convolutional neural network subunit, the first convolutional feature is modified to remove redundant features; normalization and other processing can also be performed to obtain the modified first convolutional feature. Then, based on the pooling layer of the second first convolutional neural network subunit, feature compression is performed on the modified first convolutional feature to reduce the dimensionality of the modified first convolutional feature. For example, if the dimension of the modified first convolutional feature is 100*100, the dimension of the modified first convolutional feature is reduced to 50*50 through the pooling layer. The first feature map of the interest point to be identified can then be obtained. The current first feature map is used as the new feature image of the interest point to be identified.

[0212] Similarly, the new feature image of the interest point to be identified is input into the q-th first convolutional neural network subunit. Based on the first convolutional layer of the q-th first convolutional neural network subunit, convolution calculation is performed on the new feature image of the interest point to be identified, resulting in the first convolutional feature. This first convolutional feature is the new convolutional feature of the feature image of the interest point to be identified. Then, based on the first activation layer of the q-th first convolutional neural network subunit, the first convolutional feature is modified to remove redundant features; normalization can also be performed, etc., to obtain the modified first convolutional feature. Then, based on the pooling layer of the q-th first convolutional neural network subunit, feature compression is performed on the modified first convolutional feature, thereby reducing its dimensionality. For example, if the dimension of the modified first convolutional feature is 100*100, the pooling layer reduces the dimension to 50*50. This yields the first feature map of the interest point to be identified. This current first feature map is then used as the new feature image of the interest point to be identified. This process continues until the first preset condition is met.

[0213] The first preset condition can be "determining that the value of q is equal to Q". Alternatively, the first preset condition can be receiving a first stop instruction, which instructs the generation of the first feature map of the interest point to be identified to be stopped. For example, the value of q can be 3.

[0214] The feature image obtained when the first preset condition is met is used as the intermediate feature of the interest point to be identified. Therefore, based on the feature image of the interest point to be identified by the first convolutional neural network subunit, feature extraction and feature compression processes are performed. The resulting processed feature image of the interest point to be identified can be better combined with the relevant features of the interest point to be identified.

[0215] Each of the above pooling layers is a max pooling layer, and edge features are extracted based on the max pooling layer, rather than non-texture features.

[0216] Then, after obtaining the intermediate features of the interest point to be identified, these features are input into the first second convolutional neural network subunit. Based on the second convolutional layer of the first second convolutional neural network subunit, convolution calculations are performed on the intermediate features of the interest point to be identified, resulting in the second convolutional feature. This second convolutional feature is the convolutional feature of the intermediate features of the interest point to be identified. Then, based on the second activation layer of the first second convolutional neural network subunit, the second convolutional feature is modified to remove redundant features; normalization may also be performed, etc., to obtain the modified second convolutional feature. This modified second convolutional feature is then determined as the second feature map of the interest point to be identified. Finally, the current second feature map is used as the new intermediate feature of the interest point to be identified.

[0217] The new intermediate features of the interest point to be identified are input into the second convolutional neural network subunit. Based on the second convolutional layer of the second convolutional neural network subunit, the intermediate features of the interest point to be identified are processed by convolution to obtain the second convolutional feature. It can be seen that the second convolutional feature is the convolutional feature of the intermediate features of the interest point to be identified. Then, based on the second activation layer of the second convolutional neural network subunit, the second convolutional feature is modified to remove redundant features; normalization processing can also be performed, etc., to obtain the modified second convolutional feature. The modified second convolutional feature is determined as the second feature map of the interest point to be identified. Then, the current second feature map is used as the new intermediate feature of the interest point to be identified.

[0218] Similarly, the new intermediate features of the interest point to be identified are input into the k-th second convolutional neural network subunit. Based on the second convolutional layer of the k-th second convolutional neural network subunit, the intermediate features of the interest point to be identified are processed by convolution calculation to obtain the second convolutional feature. It can be seen that the second convolutional feature is the convolutional feature of the intermediate features of the interest point to be identified. Then, based on the second activation layer of the k-th second convolutional neural network subunit, the second convolutional feature is modified to remove redundant features; normalization processing can also be performed, etc., to obtain the modified second convolutional feature. The modified second convolutional feature is determined as the second feature map of the interest point to be identified. Then, the current second feature map is used as the new intermediate feature of the interest point to be identified. This process is repeated until the second preset condition is met.

[0219] The second preset condition can be "determining that the value of k is equal to MQ". Alternatively, the second preset condition can be receiving a second stop instruction, which instructs the generation of the second feature map of the interest point to be identified to stop. For example, the value of M can be 5, the value of Q can be 3, and then MQ is 2.

[0220] The intermediate features obtained when the second preset condition is met are used as the processed feature image of the interest points to be identified.

[0221] In the above process, the pooling layer's role is to compress features; each first convolutional neural network subunit includes a pooling layer, but each second convolutional neural network subunit does not. This avoids prematurely entering the nonlinear combination stage of strong features and prevents feature loss.

[0222] In the above process, the number of convolutional kernels in each first convolutional neural network subunit is less than the preset number, and the number of convolutional kernels in each second convolutional neural network subunit is less than the preset number, thereby reducing the computational complexity of the recognition model and speeding up the recognition process.

[0223] Furthermore, the training process of the recognition model is implemented based on the dropout mechanism in neural networks.

[0224] S605. Based on the recognition model, the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period are processed to obtain the traffic information of the point of interest to be identified in the current time period.

[0225] For example, this step is the same as step S404, and will not be repeated here. However, in this step, the "feature image of the interest point to be identified" is the "processed feature image of the interest point to be identified" output by step S604.

[0226] S606. Generate a navigation path based on the traffic information of the point of interest to be identified in the current time period; control the autonomous vehicle to drive according to the navigation path.

[0227] For example, after obtaining the traffic information of the point of interest to be identified in the current time period, a navigation path can be generated using the Lu generation algorithm based on the preset starting point and the traffic information of the point of interest to be identified in the current time period. The generated navigation path is then displayed.

[0228] Then, based on the navigation path, the autonomous vehicle is controlled to drive automatically. This allows for accurate and automatic navigation based on current traffic information, precisely guiding the user or vehicle to the point of interest, enabling them to enter the point of interest promptly.

[0229] In this embodiment, based on the above embodiments, the feature image of the interest point to be identified can be processed. At least one convolutional neural network unit in the recognition model processes the feature image of the interest point to be identified, resulting in a processed feature image of the interest point. This completes the processes of feature extraction and feature compression. The resulting processed feature image of the interest point can be better combined with the relevance features of the interest point. Furthermore, during the processing, some convolutional neural network units in at least one unit have pooling layers, while others do not, thus avoiding premature entry into the nonlinear combination stage of strong features and preventing feature loss.

[0230] Figure 8 This is a schematic diagram based on the fourth embodiment of the present disclosure, as shown below. Figure 8 As shown, the model training method for interest point information recognition provided in this embodiment includes:

[0231] S801. Obtain a set of interest points to be trained; wherein, the set of interest points to be trained includes multiple interest points to be trained, and each interest point to be trained has access information and trajectory information. Access information is information that can indicate the accessibility of the interest point, and trajectory information represents the trajectory with the interest point as the key position.

[0232] For example, the execution subject in this embodiment may be a server, a terminal device, an electronic device, an autonomous vehicle, a controller on an autonomous vehicle, a remote device, a model training device or equipment for identifying points of interest, or other devices or equipment capable of executing the method of this embodiment. This embodiment is described with an electronic device as the execution subject.

[0233] Obtain a set of interest points to be trained, which includes multiple interest points. This set can be obtained from other devices or locally.

[0234] In this embodiment, the initial model needs to be trained based on the accessibility and trajectory information of the points of interest to be trained. Therefore, each point of interest to be trained needs to have accessibility and trajectory information.

[0235] Accessibility information refers to information that indicates the accessibility of a point of interest (POI). Accessibility information includes one or more of the following: accessible entrances to the POI, the orientation of the accessible entrances to the POI, the location of the POI, the opening hours of the POI, etc.

[0236] For example, for a point of interest A, the accessible passage of the point of interest refers to whether the access information of point of interest A is accessed through point a. Point a is the entrance or exit of point of interest A.

[0237] For example, for a point of interest A, the orientation of the accessible passageway of the point of interest refers to the orientation of the entrance of the point of interest A, or the orientation of the exit of the point of interest A.

[0238] For example, for a point of interest A, the location of the point of interest refers to the specific location of the entrance to the point of interest, and the specific location of the exit to the point of interest, based on the access information of the point of interest A.

[0239] For example, for a point of interest A, the opening time of the point of interest refers to the access information of point of interest A, which indicates that the point of interest is open from time 1 to time 2.

[0240] Trajectory information refers to a trajectory with points of interest (POIs) as key locations. A trajectory can be a user's driving path or a vehicle's driving path. Key locations refer to POIs as either the endpoint of the trajectory or important points along the route.

[0241] For example, for a trajectory G, trajectory G ends at point of interest A. Or, for a trajectory G, an important point along the way is point of interest A.

[0242] S802. Based on the traffic information and trajectory information of the interest points to be trained, determine the correlation features of the interest points to be trained; wherein, the correlation features characterize the correlation between the traffic information and trajectory information of the same interest point in the same time period.

[0243] For example, each point of interest to be trained has access information and trajectory information. Then, for each point of interest to be trained, the access information and trajectory information are analyzed and processed to determine the correlation information between the access information and trajectory information of that point of interest within the same time period, thereby obtaining the correlation characteristics of each point of interest within each same time period. The correlation characteristics indicate the correlation and association relationships between the access information and trajectory information of the point of interest to be trained within the same time period.

[0244] The term "period" as used above refers to a period of time; for example, a "period" is P hours, or a "period" is P days, where P is a positive integer greater than or equal to 1.

[0245] Furthermore, the aforementioned "same time period" refers to a single time period consisting of one moment and another. If, for two time periods, the length of one time period equals the length of the other, and the two moments preceding and following each of the two time periods are the same as the two moments preceding and following each of the other time period, then these two time periods are considered the same time period. For example, 13:00 to 14:00 on January 1st constitutes time period 1; 13:00 to 14:00 on January 2nd constitutes time period 2. Time period 1 and time period 2 have the same length, and since both time period 1 and time period 2 are from 13:00 to 14:00, time period 1 and time period 2 can be considered the same time period.

[0246] Alternatively, the aforementioned "same time period" refers to a time period consisting of one moment and another. If, for two time periods, the date of one time period differs from the date of the other, then these two time periods are not the same. In this case, no two time periods are identical; each time period is distinct. For example, 13:00 to 14:00 on January 1st constitutes time period 1; 13:00 to 14:00 on January 2nd constitutes time period 2. Although time period 1 and time period 2 have the same duration—time period 1 is from 13:00 to 14:00, and time period 2 is also from 13:00 to 14:00—time period 1 is a time period following January 1st, and time period 2 is a time period following January 2nd. Time period 1 and time period 2 cannot be considered the same time period; they are two distinct and independent time periods.

[0247] For example, the set of interest points to be trained includes interest point A. For interest point A, it has access information and trajectory information; these are the information of interest point A within the same time period a. Analyzing the access information and trajectory information of interest point A within the same time period a yields a correlation feature. This correlation feature indicates the correlation and association between the access information and trajectory information of interest point A within the same time period a.

[0248] S803. Train the relevance features of the points of interest to be trained to obtain the recognition model; wherein, the recognition model is used to identify the access information of the points of interest to be identified.

[0249] For example, after obtaining the relevance features of the interest points to be trained, the initial model is trained based on the relevance features of the interest points to be trained, and then the recognition model can be obtained.

[0250] The obtained recognition model can be used to identify the traffic information of the point of interest to be identified; the trajectory information of the point of interest to be identified in the current time period is input into the recognition model; then, the trajectory information of the point of interest to be identified in the current time period is processed based on the recognition model. Since the recognition model can determine the correlation between the traffic information and trajectory information of the point of interest to be identified in the same time period, the correlation between the traffic information and trajectory information of the point of interest to be identified in the current time period can be determined based on the recognition model, thereby outputting the traffic information of the point of interest to be identified in the current time period.

[0251] In one example, during training, a set of interest points to be trained can be obtained, which includes multiple interest points. Then, for each interest point, its relevance features are extracted, and the initial model is trained based on these features. The parameters of the initial model are then updated based on these relevance features. This process is repeated for the next interest point, extracting its relevance features and training the initial model based on these features. This process continues until all interest points have been processed, resulting in the recognition model.

[0252] In another example, during training, a set of interest points to be trained can be obtained, which includes multiple interest points. Then, for each interest point, relevance features are extracted. Next, for a given interest point, the initial model is trained based on its relevance features, and the parameters of the initial model are updated accordingly. This process is repeated for the next interest point until all interest points have been processed, resulting in a recognition model.

[0253] In this embodiment, multiple points of interest (POIs) to be trained are acquired. These POIs possess both accessibility information and trajectory information. Accessibility information indicates the accessibility of the POIs, while trajectory information represents the trajectory with the POIs as key locations. Based on the accessibility and trajectory information of the POIs, relevance features are determined. These relevance features indicate the correlation and association between the accessibility and trajectory information of the same POI within the same time period. The initial model is then trained based on these relevance features to obtain a recognition model. This process trains the model based on the correlation and association between the accessibility and trajectory information of the POIs, enabling the model to learn these relationships. The resulting recognition model can automatically determine the accessibility information of the POIs to be identified, resulting in an accurate and rapid model for determining the accessibility information of the POIs. Furthermore, it eliminates the need for a data collection vehicle or manual labor to acquire the accessibility information of the POIs, allowing for accurate and rapid determination of the accessibility information of the POIs to be identified; this reduces time and labor costs and improves efficiency. Furthermore, for a point of interest to be identified, only the next trajectory information of the point of interest in the current time period is needed to determine the passage information of the point of interest in the current time period. Multiple trajectory information is not required to obtain the passage information in the current time period; the passage information of the point of interest to be identified in the current time period can be determined in a timely manner.

[0254] To help readers gain a deeper understanding of the implementation principles of this disclosure, the following will be discussed in conjunction with... Figures 9-12 right Figure 8 The illustrated embodiments are further refined.

[0255] Figure 9 This is a schematic diagram based on the fifth embodiment of the present disclosure, as shown below. Figure 9 As shown, the model training method for interest point information recognition provided in this embodiment includes:

[0256] S901. Obtain user voice based on intelligent question answering; wherein, the content of the user voice includes access information for the points of interest to be trained. Access information is information that can indicate the accessibility of the points of interest.

[0257] In one example, step S901 includes the following process: obtaining the user's intent, and generating question speech based on the user's intent and the preset question text corresponding to the point of interest to be trained; sending the question speech to the user, and receiving user feedback speech.

[0258] In one example, access information includes one or more of the following: accessible passageway, the orientation of the accessible passageway, the location of the point of interest, and the opening hours of the point of interest.

[0259] For example, the execution subject in this embodiment may be a server, a terminal device, an electronic device, an autonomous vehicle, a controller on an autonomous vehicle, a remote device, a model training device or equipment for identifying points of interest, or other devices or equipment capable of executing the method of this embodiment. This embodiment is described with an electronic device as the execution subject.

[0260] First, a set of interest points to be trained needs to be obtained, which includes multiple interest points to be trained; each interest point to be trained needs to have access information and trajectory information.

[0261] Accessibility information refers to information that indicates the accessibility of a point of interest (POI). Accessibility information includes one or more of the following: accessible entrances to the POI, the orientation of the accessible entrances to the POI, the location of the POI, the opening hours of the POI, etc.

[0262] The traffic information of the interest point to be trained in this embodiment at each time period can include multiple types of information. This allows for the training of a recognition model to identify multiple types of traffic information.

[0263] Trajectory information refers to a trajectory with points of interest (POIs) as key locations. A trajectory can be a user's driving path or a vehicle's driving path. Key locations refer to POIs as either the endpoint of the trajectory or important points along the route.

[0264] In this embodiment, for each interest point in the set of interest points to be trained, the access information of each interest point to be trained can be obtained based on intelligent question answering.

[0265] If the actual traffic information of the point of interest to be trained in each time period has not been obtained before, or if the location of the point of interest to be trained is relatively marginal and it is difficult to determine the traffic information of the point of interest to be trained through multiple trajectories, the traffic information of the point of interest to be trained can be obtained through intelligent question answering based on voice interaction.

[0266] When acquiring access information for each point of interest to be trained using intelligent question answering, multiple texts are pre-set, each containing text information about a question posed to the point of interest. Based on a Text-to-Speech (TTS) model, text is extracted from the pre-set texts; then, based on the TTS model, the extracted text is used to generate the spoken question.

[0267] The generated question message is then sent to the user. For example, a phone call is made to the user, and then the question message is sent.

[0268] This allows users to reply with intelligent voice questions, and in turn, the system can obtain the user's voice messages, which contain access information about points of interest to be trained.

[0269] For example, based on a TTS (Text-to-Speech) model, text is converted into human-like spoken questions. A virtual digital human then sends these spoken questions to the user via telephone. Examples of spoken questions include: "Is there parking available in front of the store?", "Are there any parking lots nearby?", "Does the scenic area have traffic restrictions during holidays?", "Does your store face Road A?", and so on. The user can then respond with spoken responses. This allows the acquisition of the user's spoken responses, such as: "Parking is available in front of the store", "There are no parking lots nearby", "The scenic area has traffic restrictions during holidays", "The store faces Road A", and so on.

[0270] In one example, step S901 can be implemented as follows: Multiple preset question texts are pre-set, each containing text information about a question posed to the point of interest to be trained. Based on the TTS model, the preset question texts are extracted from the preset texts; based on the TTS model, question speech is generated according to the extracted preset question texts.

[0271] In intelligent question-and-answer interactions with users, to achieve smooth dialogue and effectively reduce the probability and rate of users hanging up, user intent can be obtained. Users can issue a voice message upon first speaking, or after receiving the first voice question. Then, based on the user's voice message, intent recognition methods are used to identify the user's intent.

[0272] Then, a new question speech is generated based on the user's intent. At this point, it is necessary to determine the pre-defined question text that corresponds to both the current training point of interest and the user intent, based on the traffic information to be obtained. For example, based on the correspondence between the training point of interest, the user intent, and the pre-defined question text, a pre-defined question text is determined. Then, a new question speech is generated based on the determined pre-defined question text. Alternatively, it is necessary to determine the pre-defined question text that corresponds to the current training point of interest, based on the traffic information to be obtained. For example, based on the correspondence between the training point of interest, the pre-defined question text, and the user intent, and the pre-defined question text corresponding to the current training point of interest, a complete and fluent question speech is generated, thus obtaining a new question speech.

[0273] A new voice message with a question is sent via telephone. Since the user will provide feedback based on this new voice message, the user's voice messages can be obtained.

[0274] Through the aforementioned intelligent voice interaction process, questions can be uttered to the user based on their intent, inquiring about the accessibility information of the points of interest to be trained. This enables smooth dialogue with the user and effectively reduces the probability and rate of users hanging up; thus, the accessibility information of the points of interest to be trained can be determined effectively, accurately, and quickly.

[0275] For example, based on a preset question text, a question voice is generated: "Is this location A?"; location A is a point of interest to be trained. Using an intelligent digital human, the generated question voice is sent to the user via telephone. The user's voice is received: "Yes, who is this?". At this point, it is necessary to determine the user's intent to generate a suitable new question voice to reduce the probability and proportion of the user hanging up. Based on the user's voice "Yes, who is this?", the user intent can be extracted as "The user has questions about the current call content." Then, based on the user intent and the preset question text, a new question voice is generated: "I am intelligent customer service B. Is entrance A of the store currently open?" Then, using the intelligent digital human, the generated new question voice is sent to the user via telephone. When the user responds with the new voice "Entrance A is open," the access information of the point of interest to be trained, "Entrance A is currently open," can be extracted.

[0276] S902. Determine the passage information corresponding to the user's voice based on the user's voice.

[0277] In one example, step S902 includes the following steps:

[0278] The first step of step S902 is to determine the similarity between the user's voice and the candidate voices in a preset candidate voice set, and obtain the similarity information of the candidate voices, wherein the preset candidate voice set includes at least one candidate voice.

[0279] The second step of step S902 is to determine the candidate speech with the highest similarity value to correct the speech.

[0280] The third step of step S902 is to determine the access information in the corrected speech, which is the access information corresponding to the user's speech.

[0281] In one example, the first step of step S902 includes the following process: extracting first audio information of the user's speech and extracting second audio information of the candidate speech; determining the similarity between the first audio information and the second audio information to obtain similarity information of the candidate speech.

[0282] For example, since the content of the user's speech includes the access information of the interest points to be trained, the access information of the interest points to be trained can be extracted from the user's speech.

[0283] This method uses intelligent question-and-answer to acquire user speech containing traffic information for points of interest to be trained, and then extracts the traffic information for these points from the user speech. This allows for the rapid acquisition of traffic information for many points of interest at any given time period, eliminating the need for data collection vehicles or manual methods. This facilitates the rapid training of the recognition model.

[0284] In one example, when extracting access information corresponding to a user's speech—that is, extracting access information from the user's speech—the extracted access information may be inaccurate due to the influence of the user's accent. The extracted access information may not reflect the true meaning conveyed by the user's speech. Therefore, it is necessary to correct the user's accent.

[0285] A preset candidate speech set is provided, which includes at least one candidate speech, which can be in Mandarin pronunciation; the content of the candidate speech is traffic information. This preset candidate speech set needs to be extracted first.

[0286] Alternatively, multiple preset candidate speech sets can be established, each containing at least one candidate speech, the content of which is travel information. Each candidate speech set corresponds to a geographical region. For each candidate speech set, the candidate speech within that set is spoken in the geographical region corresponding to that set. Thus, for the current user's speech, the geographical region of the user is known in advance; consequently, the preset candidate speech set matching the user's speech can be extracted.

[0287] For example, candidate voice information can include: location information, orientation information, whether parking is available, road name, whether the entrance is open, whether the exit is open, the name of the point of interest, the names of adjacent points of interest, and so on.

[0288] Then, the similarity between the user's voice and each candidate voice in the extracted preset candidate voice set is calculated to obtain the similarity information of each candidate voice.

[0289] In one example, during the process of obtaining similarity information for each candidate speech, if speech recognition and text processing are performed on the user's speech to obtain the corresponding text, and then the similarity between the text corresponding to the user's speech and the text corresponding to each candidate speech is calculated, this method is inaccurate. This is because if the user's speech has a regional accent, the text obtained from speech recognition and text processing will be inaccurate. Therefore, this method cannot be used to determine similarity information.

[0290] In this embodiment, during the process of obtaining the similarity information of each candidate speech, the similarity between the user's speech and each candidate speech in the extracted preset candidate speech set can be determined based on the audio.

[0291] One approach involves: performing audio analysis on the user's speech to obtain the first audio information; then performing audio analysis on each candidate speech to obtain the second audio information. For each candidate speech, the similarity between the first audio information of the user's speech and the second audio information of the candidate speech is calculated, yielding the similarity information cos=(X1*X2) / (norm(X1)*norm(X2)). Here, X1 is the first audio information of the user's speech, X2 is the second audio information of the candidate speech, and norm() is a function that assigns length and size to vectors in a vector space.

[0292] Another approach involves performing audio analysis on the user's speech to obtain the first audio information of the user's speech; then performing audio analysis on each candidate speech to obtain the second audio information of each candidate speech. For each candidate speech, the first audio information of the user's speech and the second audio information of the candidate speech are input into a preset binary classification model; then, based on the preset binary classification model, the audio similarity between the first audio information of the user's speech and the second audio information of the candidate speech is calculated to obtain the similarity information of the candidate speech.

[0293] By calculating the audio similarity between the first audio information of the user's speech and the second audio information of each candidate speech, the similarity information of each candidate speech can be accurately determined. This is beneficial for subsequently identifying the candidate speech most similar to the user's speech, and thus accurately correcting the user's speech.

[0294] After obtaining the similarity between the user's voice and each candidate voice in the extracted preset candidate voice set, the candidate voice with the highest similarity value is determined as the corrected voice. It can be seen that since the content of the candidate voice has common information, and the corrected voice is a candidate voice in the extracted preset candidate voice set, the corrected voice has common information.

[0295] Then, the passage information in the corrected speech can be identified as the passage information of the user's speech. This allows for the accurate extraction of the passage information from the user's speech.

[0296] For each point of interest to be trained, the user voice corresponding to the point of interest to be trained obtained based on the intelligent question answering method can be used in step S902 to determine the passage information in the user voice, and then extract the passage information of the point of interest to be trained.

[0297] Furthermore, for each point of interest to be trained in each time period, the user's voice for that point of interest in each time period is obtained based on intelligent question answering. Then, based on step S902, the traffic information in the user's voice for that point of interest in each time period is determined.

[0298] S903. Obtain the trajectory information of the interest points to be trained corresponding to the user's voice, so as to obtain the interest points to be trained in the set of interest points to be trained.

[0299] The set of interest points to be trained includes multiple interest points to be trained. Each interest point to be trained has access information and trajectory information. Access information is information that can indicate the accessibility of the interest point, and trajectory information represents the trajectory with the interest point as the key location.

[0300] In one example, the point of interest to be trained has traffic information for each time period within a preset time period set, and trajectory information for each time period within the preset time period set; wherein the preset time period set includes at least one time period.

[0301] The points of interest to be trained have real traffic information, which is the traffic information of the last time period within the preset time period set.

[0302] For example, trajectory information of the points of interest to be trained can be obtained from other devices or databases.

[0303] Trajectory information refers to a trajectory with points of interest (POIs) as key locations. A trajectory can be a user's driving path or a vehicle's driving path. Key locations refer to POIs as either the endpoint of the trajectory or important points along the route.

[0304] Since the access information and trajectory information of each point of interest to be trained are obtained, various information about that point of interest can be obtained. Multiple points of interest to be trained constitute a set of points of interest to be trained.

[0305] In one example, since an initial model needs to be trained to obtain a recognition model for identifying traffic information of points of interest to be identified, it is necessary to obtain the traffic information and trajectory information of the points of interest to be trained in each time period. Each time period constitutes a preset time period set. Therefore, each point of interest to be trained has traffic information and trajectory information for each time period within the preset time period set.

[0306] It can be seen that, for each point of interest to be trained, the traffic information of the point of interest to be trained in each time period is the actual traffic information of the point of interest to be trained in each time period. Furthermore, for each point of interest to be trained, the trajectory information of the point of interest to be trained in each time period is the actual trajectory information of the point of interest to be trained in each time period.

[0307] In this system, the access information for each trainable interest point (interest point) within the set of training interest points, across different time periods within a preset time period set, contains both positive and negative samples. Different weights can be assigned to the positive and negative samples for each trainable interest point, ensuring a balanced information flow across all time periods. For example, access information represents whether the entry point for a trainable interest point is open. Therefore, for each trainable interest point, if its access information in one time period within the preset time period set indicates that the entry point is open (a positive sample), and if its access information in the next time period within the preset time period set indicates that the entry point is not open (a negative sample), then...

[0308] For example, the set of interest points to be trained includes two interest points, namely POI1 and POI2. The preset time period set has three time periods, namely time period 1, time period 2, and time period 3.

[0309] The passage information for the training point of interest (POI1) in time period 1 indicates that the entrance to POI1 is open during time period 1; and POI1 has trajectory information within time period 1. The passage information for the training point of interest (POI1) in time period 2 indicates that the entrance to POI1 is not open during time period 2; and POI1 has trajectory information within time period 2. The passage information for the training point of interest (POI1) in time period 3 indicates that the entrance to POI1 is open during time period 3; and POI1 has trajectory information within time period 3.

[0310] The access information for POI2 under training time period 1 indicates that the entrance to POI2 is not open during time period 1; and POI2 has trajectory information within time period 1. The access information for POI2 under training time period 2 indicates that the entrance to POI2 is open during time period 2; and POI2 has trajectory information within time period 2. The access information for POI2 under training time period 3 indicates that the entrance to POI2 is not open during time period 3; and POI2 has trajectory information within time period 3.

[0311] S904. Based on the traffic information and trajectory information of the interest points to be trained, determine the correlation features of the interest points to be trained; wherein, the correlation features characterize the correlation between the traffic information and trajectory information of the same interest point in the same time period.

[0312] In one example, the interest points to be trained have relevance features for each time period within a preset time period set; wherein the preset time period set includes at least one time period.

[0313] For example, each point of interest to be trained has access information and trajectory information. Then, for each point of interest to be trained, the access information and trajectory information are analyzed and processed to determine the correlation information between the access information and trajectory information of that point of interest within the same time period, thereby obtaining the correlation characteristics of each point of interest within each same time period. The correlation characteristics indicate the correlation and association relationships between the access information and trajectory information of the point of interest to be trained within the same time period.

[0314] The term "time period" as used above refers to a period of time; for example, a "time period" is P hours, or a "time period" is P days, where P is a positive integer greater than or equal to 1. For an explanation of "time period," please refer to step S802, which will not be repeated here.

[0315] Since each point of interest to be trained possesses both traffic information and trajectory information for each time period within a preset time period set, the relevance features of the point of interest to be trained can be determined based on its traffic and trajectory information for each time period. Therefore, it can be concluded that each point of interest to be trained possesses relevance features for each time period within the preset time period set.

[0316] For example, for a point of interest A to be trained, this embodiment can be used to obtain the traffic information and trajectory information of point of interest A in time period 1. Then, based on the traffic information and trajectory information of point of interest A in time period 1, a correlation feature 1 for point of interest A in time period 1 can be determined. Correlation feature 1 characterizes the correlation and association between the traffic information and trajectory information of point of interest A in time period 1. Similarly, this embodiment can be used to obtain the traffic information and trajectory information of point of interest A in time period 2. Then, based on the traffic information and trajectory information of point of interest A in time period 2, a correlation feature 2 for point of interest A in time period 2 can be determined. Correlation feature 2 characterizes the correlation and association between the traffic information and trajectory information of point of interest A in time period 2. This process continues, resulting in the correlation features of the point of interest to be trained in each time period within a preset time period set.

[0317] S905. Input the relevance features of the interest points to be trained into the initial model to obtain the predicted traffic information of the interest points to be trained; wherein, the predicted traffic information is the traffic information of the interest points to be trained in the last time period.

[0318] In one example, step S905 includes the following steps:

[0319] The first step of step S905 is to obtain the feature image of the interest point to be trained; wherein, the feature image is used to indicate the features of the interest point.

[0320] The second step of step S905 is to input the feature image of the interest point to be trained and the correlation features of the interest point to be trained into the initial model to obtain the prediction information of the interest point to be trained.

[0321] For example, the relevance features of the interest points to be trained are input into the initial model, and then the relevance features of the interest points to be trained are processed based on the initial model to obtain the common information for the prediction of the interest points to be trained.

[0322] For each point of interest to be trained, according to step S904, the relevance features of the point of interest to be trained under each time period within the preset time period set are obtained. For each point of interest to be trained, the relevance features of the point of interest to be trained under each time period within the preset time period set can be input into the initial model for processing, thereby enabling the initial model to learn the relevance features. Furthermore, based on the initial model, the traffic information of the point of interest to be trained under the last time period within the preset time period set is output, that is, the predicted traffic information of the point of interest to be trained is obtained.

[0323] In one example, when implementing step S905, the initial model can be trained based on the feature image of the interest point to be trained and its correlation features. The feature image of the interest point to be trained can indicate the features of the interest point. This allows for the incorporation of more features to train the initial model.

[0324] First, feature images of the interest points to be trained need to be obtained. These feature images are generated based on the interest point information of the interest points. This interest point information includes one or more of the following: the specific location of the interest point, the road network conditions corresponding to the interest point, the driving trajectory corresponding to the interest point, etc. Therefore, the feature images can indicate the features of the interest points to be trained.

[0325] Then, for each point of interest to be trained, the feature image of the point of interest to be trained and the relevance features of the point of interest to be trained can be input into the initial model to train the initial model. During the training process, the traffic information of the point of interest to be trained in the last time period within the preset time period set is predicted to obtain the predicted traffic information of the point of interest to be trained.

[0326] Therefore, by incorporating more features to train the initial model, the initial model can be trained to achieve higher accuracy. The resulting recognition model has better accuracy and can more accurately identify the traffic information of the points of interest to be identified.

[0327] In one example, the first step of S905 includes the following process:

[0328] Step 1: Obtain the interest point information of the interest points to be trained; wherein, the interest point information includes at least one feature information.

[0329] Step 2: If at least one feature includes road network information and trajectory data information; wherein the road network information is the road network information of the location of the point of interest; and the trajectory data information is the trajectory data information with the point of interest as the key location, then the trajectory data information is subjected to deviation correction processing based on the road network information to obtain the corrected trajectory data information.

[0330] Step 3: Based on the feature information in the interest point information of the interest point to be trained from the data source, perform data fusion to obtain the feature image of the interest point to be trained.

[0331] In one example, at least one feature includes one or more of the following: location information, map information, road network information, trajectory data information, and areal area information.

[0332] Among them, location information is the location information of the point of interest; map information is the map information of the location of the point of interest; road network information is the road network information of the location of the point of interest; trajectory data information is the trajectory data information with the point of interest as the key location; and areal region information is the areal region information of the point of interest.

[0333] In one example, the second step of S905 includes the following process: inputting the feature image of the interest point to be trained and the correlation features of the interest point to be trained into the first fully connected layer in the initial model for feature fusion to obtain the fused features of the interest point to be trained; inputting the fused features of the interest point to be trained into the second fully connected layer in the initial model for prediction processing to obtain the predicted access information of the interest point to be trained.

[0334] For example, the following process can be used to implement step S905.

[0335] First, it is necessary to obtain the feature image of the interest point to be trained. The feature image is generated based on the interest point information of the interest point to be trained; the feature image can indicate the features of the interest point to be trained. Therefore, it is necessary to obtain the interest point information of the interest point to be trained so that the feature image of the interest point to be trained can be generated based on the interest point information.

[0336] For each point of interest to be trained, the information of the point of interest includes at least one feature. The feature can be any one of the following: location information, map information, road network information, trajectory data information, and areal region information.

[0337] Location information refers to the location information of a point of interest; location information can be the specific location of the area, road, or intersection where the point of interest is located, or it can be the latitude and longitude of the point of interest.

[0338] Map information refers to the map information of the location of the point of interest; map information includes detailed map content matching the point of interest.

[0339] Road network information refers to the road network information of the location of a point of interest; road network information can be obtained from the road network system.

[0340] Trajectory data refers to trajectory data with points of interest (POIs) as key locations. A key location refers to a POI being either the endpoint of the trajectory or an important point along the trajectory. Trajectory data can be a user's driving trajectory or a vehicle's driving trajectory.

[0341] Area information refers to the area information of points of interest. Area information includes one or more of the following: Area of ​​Interest (AOI) area information and Building Unknown (BUD) area information. The "Area of ​​Interest" can also be called the "Area of ​​Interest".

[0342] The above process yields multiple feature information, ensuring feature completeness for the interest points to be trained. Then, at least one feature from the interest point information is integrated into the relevance features of the interest point to train the model. This allows the resulting recognition model to learn more information about the interest points.

[0343] For example, Figure 10 This is an illustration based on the feature information provided in this disclosure. Figure 2 ,like Figure 10 As shown, the set of interest points to be trained includes three interest points: POI1, POI2, and POI3. For POI1, feature maps 1001a (BUD areal region information), 1002a (AOI areal region information), 1003a (current interest point information), 1004a (road network information), 1005a (trajectory link), and 1006a (trajectory data information) can be obtained. The trajectory link refers to the corrected trajectory data obtained by performing deviation correction processing on the trajectory data based on the road network information. The current interest point information includes the location information, name, etc., of the interest point. The pixel parameters of the feature map for each feature information can be 60*60; the size of the physical world covered by the feature map for each feature information is 80 meters * 80 meters.

[0344] For the point of interest to be trained (POI1), the above-mentioned feature information of the point of interest to be trained (POI1) is fused using the method of this embodiment to obtain the feature image 1007a of the point of interest to be trained (POI1).

[0345] For the point of interest (POI2) to be trained, the following feature maps can be obtained: BUD area information of the point of interest (POI2) to be trained (1001b), BOI area information of the point of interest (POI2) to be trained (1002b), current information of the point of interest (POI2) to be trained (1003b), road network information of the point of interest (POI2) to be trained (1004b), trajectory link of the point of interest (POI2) to be trained (1005b), and trajectory data information of the point of interest (POI2) to be trained (1006b).

[0346] For the point of interest POI2 to be trained, the above-mentioned feature information of the point of interest POI2 to be trained is fused in the manner of this embodiment to obtain the feature image 1007b of the point of interest POI2 to be trained.

[0347] For the point of interest (POI3) to be trained, the following feature maps can be obtained: BUD area information 1001c, COI area information 1002c, current information of the point of interest 1003c, road network information 1004c, trajectory link information 1005c, and trajectory data information 1006c.

[0348] For the point of interest POI3 to be trained, the above-mentioned feature information of the point of interest POI3 to be trained is fused in the manner of this embodiment to obtain the feature image 1007c of the point of interest POI3 to be trained.

[0349] Furthermore, based on the feature images 1007a, 1007b, and 1007c of the interest points POI1, they can be fused into a single feature map 1008.

[0350] Then, for each point of interest to be trained, after obtaining the point of interest information, if the point of interest information includes road network information and trajectory data, the trajectory data can be repaired. Road network information can characterize the real and accurate road conditions at the location of the point of interest. However, trajectory data is obtained from the user's or vehicle's driving process, and driving deviations may occur during the user's or vehicle's driving process, thus the trajectory data may be inaccurate, or it may include redundant information.

[0351] Therefore, deviation correction processing can be performed on trajectory data based on road network information. In one example, a Hidden Markov Model (HMM) can be used to process the trajectory data based on road network information. This corrects the inherent deviation between the road network information and the trajectory data, making the trajectory data closer to the road network information, and then outputting the corrected trajectory data. This deviation correction processing based on road network information avoids the introduction of noise into the feature image due to trajectory data offset.

[0352] Next, the feature information from the interest point information of the training interest points needs to be fused to obtain the feature image of the training interest points. This fusion can be performed using a data source-based approach, rather than using the RGB (Red, Green, Blue) method of the image.

[0353] For each point of interest to be trained, data fusion is performed on the feature information in the point of interest information based on the data source. This allows the feature information to be fused relatively independently, without one feature information interfering with another. This results in the information in the feature image of the point of interest to be trained having the characteristic of lossless data representation. That is, it overcomes the data loss problem that occurs when fusing feature information based on RGB. The information in the feature image of the point of interest to be trained is accurate and the information is not interfered with.

[0354] For each interest point to be trained, a feature image of the interest point is generated based on the feature information in the interest point information; thus, the feature image of the interest point can indicate the features of the interest point.

[0355] Then, for each interest point to be trained, its feature image and related features can be input into the initial model to train it. The initial model includes two fully connected layers (FC), namely the first fully connected layer and the second fully connected layer.

[0356] During training, for each point of interest to be trained, the feature image of the point of interest to be trained and the relevance features of the point of interest to be trained in each time period within the preset time period set are input into the first fully connected layer in the initial model; based on the feature image of the point of interest to be trained and the relevance features of the point of interest to be trained in each time period within the preset time period set, feature fusion is performed on the first fully connected layer to obtain the fused features of the point of interest to be trained.

[0357] For each point of interest to be trained, the fused features of that point are input into the second fully connected layer of the initial model. Based on these fused features, traffic information is predicted. Here, the relevance features of the point of interest to be trained in each time period within a preset time period set are input into the initial model for processing. The initial model then learns these relevance features and the relationship between traffic information and trajectory information. Furthermore, when predicting traffic information based on the fused features of the point of interest to be trained, the traffic information for the last time period within the preset time period set is predicted, resulting in the predicted traffic information for the point of interest to be trained.

[0358] This approach utilizes multiple fully connected layers to fuse and predict the feature images of the points of interest to be trained, as well as their related features. This allows for better integration of these features. Furthermore, the use of multiple fully connected layers enables non-linear processing of these features.

[0359] S906. Based on the predicted traffic information of the interest points to be trained and the actual traffic information of the interest points to be trained, the parameters of the initial model are adjusted to obtain the recognition model. The recognition model is used to identify the traffic information of the interest points to be identified.

[0360] For example, step S905 can output the predicted traffic information of the interest point to be trained, which is the traffic information of the interest point to be trained in the last time period.

[0361] Since the traffic information of each point of interest to be trained in each time period is the actual traffic information of the point of interest in each time period, the point of interest to be trained has real traffic information, which is the traffic information of the point of interest in the last time period within the preset time period set.

[0362] Based on the predicted traffic information of the interest points to be trained, as well as the actual traffic information of the interest points to be trained, the parameters of the initial model can be adjusted to obtain the recognition model.

[0363] Based on the predicted and actual traffic information, the initial model is trained so that it can learn relevant features and be trained well. This allows the resulting recognition model to identify the relevant features of the points of interest to be identified, thereby obtaining the traffic information of the points of interest to be identified.

[0364] In one example, during training, based on steps S901-S903, a set of interest points to be trained is obtained, which includes multiple interest points. Then, for each interest point, based on step S904, the relevance features of that interest point are extracted. Based on steps S905-S906, the initial model is trained using these relevance features, and the parameters of the initial model are updated accordingly. Then, for the next interest point, the relevance features are extracted again based on step S904, and the initial model is trained again based on these relevance features, and the parameters of the initial model are updated accordingly. This process continues until all interest points have been processed, resulting in the recognition model.

[0365] In another example, during training, based on steps S901-S903, a set of interest points to be trained can be obtained, which includes multiple interest points. Then, for each interest point, relevance features are extracted based on step S904. Then, for a given interest point, steps S905-S906 train the initial model based on its relevance features, and update the parameters of the initial model based on those features; then, for the next interest point, steps S905-S906 train the initial model based on its relevance features, and update the parameters of the initial model based on those features; and so on, until all interest points have been processed. This yields a recognition model. The recognition model is used to identify the access information of the interest points to be identified.

[0366] In this embodiment, based on the above embodiments, user voice is acquired using an intelligent question-and-answer method, and the traffic information of the points of interest to be trained is extracted from the user voice; the trajectory information of the points of interest to be trained is also acquired; thus, each point of interest to be trained is obtained. The intelligent question-and-answer method can quickly acquire a large amount of traffic information of the points of interest to be trained, reducing time and labor costs; it is beneficial to accelerate the training process of this embodiment. Furthermore, the acquired user voice can be corrected to obtain accurate traffic information. Based on the traffic information and trajectory information of the points of interest to be trained, the relevance features of the points of interest to be trained are determined; the relevance features characterize the correlation between the traffic information and trajectory information of the same point of interest in the same time period. Based on the relevance features and feature images of the points of interest to be trained, the initial model is trained; thus, more features are added to train the initial model, resulting in higher accuracy. The obtained recognition model has good accuracy, and it can more accurately identify the traffic information of the points of interest to be identified. During training, predicted traffic information for the interest points to be trained is obtained. Then, based on this predicted traffic information and the actual traffic information of the interest points, the parameters of the initial model are adjusted to obtain the recognition model. This allows the initial model to learn relevant features, enabling effective training. The resulting recognition model can then identify the relevant features of the interest points to be identified, thus obtaining their traffic information.

[0367] Figure 11 This is a schematic diagram based on the sixth embodiment of the present disclosure, as shown below. Figure 11 As shown, the model training method for interest point information recognition provided in this embodiment includes:

[0368] S1101. Obtain user voice based on intelligent question answering; wherein, the content of the user voice includes access information for the points of interest to be trained. Access information is information that can indicate the accessibility of the points of interest.

[0369] For example, the execution subject in this embodiment may be a server, a terminal device, an electronic device, an autonomous vehicle, a controller on an autonomous vehicle, a remote device, a model training device or equipment for identifying points of interest, or other devices or equipment capable of executing the method of this embodiment. This embodiment is described with an electronic device as the execution subject.

[0370] This step is the same as step S901 and will not be repeated here.

[0371] S1102. Determine the passage information corresponding to the user's voice based on the user's voice.

[0372] For example, this step is the same as step S902, and will not be described again.

[0373] S1103. Obtain the trajectory information of the interest points to be trained corresponding to the user's voice, so as to obtain the interest points to be trained in the set of interest points to be trained.

[0374] The set of interest points to be trained includes multiple interest points to be trained. Each interest point to be trained has access information and trajectory information. Access information is information that can indicate the accessibility of the interest point, and trajectory information represents the trajectory with the interest point as the key location.

[0375] For example, this step is the same as step S903, and will not be described again.

[0376] S1104. Based on the traffic information and trajectory information of the interest points to be trained, determine the correlation features of the interest points to be trained; wherein, the correlation features characterize the correlation between the traffic information and trajectory information of the same interest point in the same time period.

[0377] For example, this step is the same as step S904, and will not be described again.

[0378] S1105. Obtain the feature image of the interest point to be trained; wherein, the feature image is used to indicate the features of the interest point.

[0379] For example, this step is the same as the first step in step S905, and will not be described again.

[0380] S1106. Input the feature image of the interest point to be trained into at least one convolutional neural network unit in the initial model for processing to obtain the processed feature image of the interest point to be trained.

[0381] In one example, at least one convolutional neural network unit includes n first convolutional neural network sub-units and mn second convolutional neural network sub-units. The structure of the first convolutional neural network sub-units differs from the structure of the second convolutional neural network sub-units. n is a positive integer greater than 1, m is a positive integer greater than 1, and m is greater than n. Step S1106 includes the following steps:

[0382] The first step of step S1106. Repeat the following steps until the first preset condition is met: Input the feature image of the interest point to be trained into the i-th first convolutional neural network sub-unit in the initial model for processing to obtain the first feature map of the interest point to be trained; determine the first feature map of the interest point to be trained as the new feature image of the interest point to be trained, and determine the value of i as i+1; where the initial value of i is 1, and i is a positive integer greater than or equal to 1 and less than or equal to n; the new feature image obtained when the first preset condition is met is the intermediate feature of the interest point to be trained.

[0383] The second step of step S1106. Repeat the following steps until the second preset condition is met: input the intermediate features of the interest point to be trained into the j-th second convolutional neural network sub-unit in the initial model for processing to obtain the second feature map of the interest point to be trained; determine the second feature map of the interest point to be trained as the new intermediate feature, and determine that the value of j is j+1; where the initial value of j is 1, and j is a positive integer greater than or equal to 1 and less than or equal to mn.

[0384] Among them, the intermediate features obtained when the second preset condition is met are the processed feature images of the interest points to be trained.

[0385] For example, for each interest point to be trained, after obtaining the feature image of the interest point to be trained, the feature image needs to be processed based on the initial model.

[0386] Figure 12 It is based on the structural diagram of the initial model provided in this disclosure, such as Figure 12 As shown, the initial model 1200 includes an input layer 1201, at least one convolutional neural network unit 1202, a long short-term memory (LSTM) network unit 1203, a first fully connected layer 1204, a second fully connected layer 1205, and an output layer 1206.

[0387] For each interest point to be trained, its feature image is input to at least one convolutional neural network unit 1202 through the input layer 1201. The feature image of the interest point is processed by the at least one convolutional neural network unit 1202 to obtain a processed feature image of the interest point, thereby completing the processes of feature extraction and feature compression. The resulting processed feature image of the interest point can be better combined with the relevant features of the interest point.

[0388] For each point of interest to be trained, the LSTM unit 1203 processes the traffic information and trajectory information of the point of interest to be trained in each time period to obtain the correlation features of the point of interest to be trained in each time period.

[0389] For example, the preset time period set includes three time periods: time period 1, time period 2, and time period 3. For the point of interest (POI1) to be trained, based on the traffic and trajectory information of POI1 in time period 1, the relevance features of POI1 in time period 1 are obtained; based on the traffic and trajectory information of POI1 in time period 2, the relevance features of POI1 in time period 2 are obtained. Then, based on the LSTM unit 1203, the relevance features of POI1 in each time period are obtained.

[0390] For each point of interest to be trained, the correlation features of the point of interest at each time period and the processed feature image of the point of interest are input into the first fully connected layer 1204 and the second fully connected layer 1205 for processing in sequence; the predicted passage information of the point of interest to be trained can be output through the output layer 1206.

[0391] In one example, the technical solution for step S1106 can be implemented in the following manner.

[0392] The initial model includes at least one convolutional neural network (CNN) unit; the number of CNN units is m, where m is a positive integer greater than 1. Furthermore, the m CNN units include n first CNN sub-units and mn second CNN sub-units, where n is a positive integer greater than 1 and less than m. The structure of the first CNN sub-units differs from that of the second CNN sub-units, and thus the feature images of the points of interest to be trained are processed based on the CNN units with different structures.

[0393] For each interest point to be trained, after acquiring its feature image, the feature image is input into the first convolutional neural network subunit. The feature image is processed by the first convolutional neural network subunit to obtain a first feature map of the interest point. This first feature map is then used as the new feature image of the interest point. The new feature image is then input into the second convolutional neural network subunit. The new feature map is processed by the second convolutional neural network subunit to obtain a new first feature map of the interest point. This new first feature map is then used as the new feature image of the interest point. This process continues until the new feature image of the interest point is input into the i-th convolutional neural network subunit. The new feature map is processed by the i-th convolutional neural network subunit to obtain a new first feature map of the interest point. The current first feature map is then used as the new feature map of the interest point, and the value of i is incremented by 1. i is a positive integer greater than or equal to 1 and less than or equal to n. This process continues until the first preset condition is met.

[0394] The first preset condition can be "determining that the value of i is equal to n". Alternatively, the first preset condition can be receiving a first stop instruction, which is used to indicate the cessation of generating the first feature map of the interest points to be trained.

[0395] The feature image obtained when the first preset condition is met is used as the intermediate feature of the interest point to be trained.

[0396] Then, for each interest point to be trained, after obtaining the intermediate features of the interest point, the intermediate features are input into the first second convolutional neural network subunit. Based on the feature image of the interest point to be trained in the first second convolutional neural network subunit, the second feature map of the interest point to be trained is obtained. Then, the current second feature map is used as the new intermediate feature of the interest point to be trained. The new intermediate features of the interest point to be trained are input into the second second convolutional neural network subunit. Based on the new feature image of the interest point to be trained in the second second convolutional neural network subunit, the second feature map of the interest point to be trained is obtained. Then, the current second feature map is used as the new intermediate feature of the interest point to be trained. And so on, the new intermediate features of the interest point to be trained are input into the j-th second convolutional neural network subunit. Based on the new feature image of the interest point to be trained in the j-th second convolutional neural network subunit, the second feature map of the interest point to be trained is obtained. Then, the current second feature map is used as the new intermediate feature of the interest point to be trained, and the value of j is incremented by 1. j is a positive integer greater than or equal to 1 and less than or equal to mn. This process continues until the second preset condition is met.

[0397] The second preset condition can be "determining that the value of j is equal to mn". Alternatively, the second preset condition can be receiving a second stop instruction, which is used to indicate the cessation of generating the second feature map of the interest points to be trained.

[0398] The intermediate features obtained when the second preset condition is met are used as the processed feature image of the interest points to be trained.

[0399] Therefore, the feature images of the points of interest to be trained are processed based on convolutional neural network units with different structures. For example, the first convolutional neural network subunit has fewer structural layers than the second convolutional neural network subunit; or, the second convolutional neural network subunit has fewer structural layers than the first convolutional neural network subunit; thereby reducing the complexity of the training process and speeding up the training process.

[0400] In one example, the i-th first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer. The first step of step S1106, "inputting the feature image of the interest point to be trained into the i-th first convolutional neural network subunit in the initial model for processing to obtain the first feature map of the interest point to be trained," includes the following process: inputting the feature image of the interest point to be trained into the first convolutional layer for convolution processing to obtain the first convolutional feature; wherein, the first convolutional feature represents the convolutional feature of the feature image of the interest point to be trained; inputting the first convolutional feature into the first activation layer for correction processing to obtain the corrected first convolutional feature; inputting the corrected first convolutional feature into the pooling layer for feature compression processing to obtain the first feature map of the interest point to be trained; wherein, the first feature map of the interest point to be trained represents the feature image of the interest point to be trained.

[0401] The j-th second convolutional neural network subunit includes a second convolutional layer and a second activation layer. The second step of step S1106, "inputting the intermediate features of the interest point to be trained into the j-th second convolutional neural network subunit in the initial model for processing to obtain the second feature map of the interest point to be trained," includes the following process: inputting the intermediate features of the interest point to be trained into the second convolutional layer for convolution processing to obtain the second convolutional features; wherein, the second convolutional features represent the convolutional features of the intermediate features of the interest point to be trained; inputting the second convolutional features into the second activation layer for correction processing to obtain the second feature map of the interest point to be trained.

[0402] For example, the technical solution of step S1106 can be implemented in the following manner.

[0403] The initial model includes at least one convolutional neural network unit; the number of at least one convolutional neural network unit is m, where m is a positive integer greater than 1. Furthermore, the m convolutional neural network units include n first convolutional neural network sub-units and mn second convolutional neural network sub-units, where n is a positive integer greater than 1 and n is less than m.

[0404] Each first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer. However, each second convolutional neural network subunit only includes a second convolutional layer and a second activation layer. Therefore, the second convolutional neural network subunit does not require a pooling layer.

[0405] For each interest point to be trained, after obtaining its feature image, the feature image is input into the first convolutional neural network subunit. Based on the first convolutional layer of the first convolutional neural network subunit, convolution calculations are performed on the feature image of the interest point to be trained, resulting in the first convolutional feature. This first convolutional feature is the convolutional feature of the feature image of the interest point to be trained. Then, based on the first activation layer of the first convolutional neural network subunit, the first convolutional feature is modified to remove redundant features; normalization may also be performed, etc., to obtain the modified first convolutional feature. Then, based on the pooling layer of the first convolutional neural network subunit, feature compression is performed on the modified first convolutional feature, thereby reducing its dimensionality. For example, if the dimension of the modified first convolutional feature is 100*100, the pooling layer reduces the dimension to 50*50. This yields the first feature map of the interest point to be trained. Use the current first feature map as the new feature image for the interest points to be trained.

[0406] The new feature image of the interest point to be trained is input into the second first convolutional neural network subunit. Based on the first convolutional layer of the second first convolutional neural network subunit, convolution calculation is performed on the new feature image of the interest point to be trained, resulting in the first convolutional feature. This first convolutional feature is the new convolutional feature of the feature image of the interest point to be trained. Then, based on the first activation layer of the second first convolutional neural network subunit, the first convolutional feature is modified to remove redundant features; normalization may also be performed, etc., to obtain the modified first convolutional feature. Then, based on the pooling layer of the second first convolutional neural network subunit, feature compression is performed on the modified first convolutional feature, thereby reducing its dimensionality. For example, if the dimension of the modified first convolutional feature is 100*100, the pooling layer reduces the dimension to 50*50. This yields the first feature map of the interest point to be trained. This current first feature map is then used as the new feature image of the interest point to be trained.

[0407] Similarly, the new feature image of the interest point to be trained is input into the i-th first convolutional neural network subunit. Based on the first convolutional layer of the i-th first convolutional neural network subunit, convolution calculation is performed on the new feature image of the interest point to be trained, resulting in the first convolutional feature. This first convolutional feature is the new convolutional feature of the feature image of the interest point to be trained. Then, based on the first activation layer of the i-th first convolutional neural network subunit, the first convolutional feature is modified to remove redundant features; normalization can also be performed, etc., to obtain the modified first convolutional feature. Then, based on the pooling layer of the i-th first convolutional neural network subunit, feature compression is performed on the modified first convolutional feature, thereby reducing its dimensionality. For example, if the dimension of the modified first convolutional feature is 100*100, the pooling layer reduces the dimension to 50*50. This yields the first feature map of the interest point to be trained. The current first feature map is then used as the new feature image of the interest point to be trained. This process continues until the first preset condition is met.

[0408] The first preset condition can be "determining that the value of i is equal to n". Alternatively, the first preset condition can be receiving a first stop instruction, which is used to indicate the cessation of generating the first feature map of the interest points to be trained. For example, the value of n can be 3.

[0409] The feature image obtained when the first preset condition is met is used as the intermediate feature of the interest point to be trained. Therefore, based on the feature image of the interest point to be trained by the first convolutional neural network subunit, feature extraction and feature compression processes are performed. The resulting processed feature image of the interest point to be trained can be better combined with the relevant features of the interest point to be trained.

[0410] Each of the above pooling layers is a max pooling layer, and edge features are extracted based on the max pooling layer, rather than non-texture features.

[0411] Then, for each interest point to be trained, after obtaining the intermediate features of the interest point, these intermediate features are input into the first second convolutional neural network subunit. Based on the second convolutional layer of the first second convolutional neural network subunit, convolution calculations are performed on the intermediate features of the interest point to be trained, resulting in the second convolutional feature. It can be seen that the second convolutional feature is the convolutional feature of the intermediate features of the interest point to be trained. Then, based on the second activation layer of the first second convolutional neural network subunit, the second convolutional feature is modified to remove redundant features; normalization may also be performed, etc., to obtain the modified second convolutional feature. This modified second convolutional feature is determined as the second feature map of the interest point to be trained. Finally, the current second feature map is used as the new intermediate feature of the interest point to be trained.

[0412] The new intermediate features of the interest point to be trained are input into the second convolutional neural network subunit. Based on the second convolutional layer of the second convolutional neural network subunit, the intermediate features of the interest point to be trained are processed by convolution to obtain the second convolutional features. It can be seen that the second convolutional features are the convolutional features of the intermediate features of the interest point to be trained. Then, based on the second activation layer of the second convolutional neural network subunit, the second convolutional features are modified to remove redundant features; normalization may also be performed, etc., to obtain the modified second convolutional features. The modified second convolutional features are determined as the second feature map of the interest point to be trained. Then, the current second feature map is used as the new intermediate features of the interest point to be trained.

[0413] Similarly, the new intermediate features of the interest point to be trained are input into the j-th second convolutional neural network subunit. Based on the second convolutional layer of the j-th second convolutional neural network subunit, the intermediate features of the interest point to be trained are processed by convolution to obtain the second convolutional feature. It can be seen that the second convolutional feature is the convolutional feature of the intermediate features of the interest point to be trained. Then, based on the second activation layer of the j-th second convolutional neural network subunit, the second convolutional feature is modified to remove redundant features; normalization may also be performed, etc., to obtain the modified second convolutional feature. The modified second convolutional feature is determined as the second feature map of the interest point to be trained. Then, the current second feature map is used as the new intermediate feature of the interest point to be trained. This process is repeated until the second preset condition is met.

[0414] The second preset condition can be "determining that the value of j is equal to mn". Alternatively, the second preset condition can be receiving a second stop instruction, which is used to indicate the cessation of generating the second feature map of the interest points to be trained. For example, if the value of m is 5 and the value of n is 3, then mn is 2.

[0415] The intermediate features obtained when the second preset condition is met are used as the processed feature image of the interest points to be trained.

[0416] In the above process, the pooling layer's role is to compress features; each first convolutional neural network subunit includes a pooling layer, but each second convolutional neural network subunit does not. This avoids prematurely entering the nonlinear combination stage of strong features and prevents feature loss.

[0417] In the above process, the number of convolutional kernels in each first convolutional neural network subunit is less than the preset number, and the number of convolutional kernels in each second convolutional neural network subunit is less than the preset number, thereby reducing the training complexity of the model and speeding up the training process.

[0418] Furthermore, the initial model training process is implemented based on the dropout mechanism in neural networks.

[0419] S1107. Input the feature image of the interest point to be trained and the correlation features of the interest point to be trained into the initial model to obtain the prediction information of the interest point to be trained.

[0420] For example, this step refers to the second step of step S905, and will not be repeated here. However, in this step, the "feature image of the interest point to be trained" is the "processed feature image of the interest point to be trained" output in step S1106.

[0421] S1108. Based on the predicted traffic information of the interest points to be trained and the actual traffic information of the interest points to be trained, the parameters of the initial model are adjusted to obtain the recognition model. The recognition model is used to identify the traffic information of the interest points to be identified.

[0422] For example, this step is the same as the second step of step S906, and will not be described again.

[0423] In this embodiment, based on the above embodiments, the feature image of the interest point to be trained can be processed. At least one convolutional neural network unit in the initial model processes the feature image of the interest point to be trained, resulting in a processed feature image of the interest point. This completes the processes of feature extraction and feature compression. The resulting processed feature image of the interest point to be trained can be better combined with the relevance features of the interest point. Furthermore, during the processing, some convolutional neural network units in at least one convolutional neural network unit have pooling layers, while others do not, thereby avoiding premature entry into the nonlinear combination stage of strong features and preventing feature loss.

[0424] Figure 13 This is a schematic diagram based on the seventh embodiment of the present disclosure, as shown below. Figure 13 As shown, the point of interest information identification device 1300 provided in this embodiment includes:

[0425] The first acquisition unit 1301 is used to acquire the initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key position.

[0426] The second acquisition unit 1302 is used to acquire the relevance features of the point of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest.

[0427] The determining unit 1303 is used to determine the access information of the point of interest to be identified in the current time period based on the trajectory information and correlation characteristics of the point of interest to be identified in the current time period.

[0428] The apparatus in this embodiment can execute the technical solutions in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.

[0429] Figure 14 This is a schematic diagram based on the eighth embodiment of the present disclosure, as shown below. Figure 14 As shown, the point of interest information identification device 1400 provided in this embodiment includes:

[0430] The first acquisition unit 1401 is used to acquire the initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period, and the trajectory information represents the trajectory with the point of interest as the key position.

[0431] The second acquisition unit 1402 is used to acquire the relevance features of the point of interest to be identified; wherein, the relevance features characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest.

[0432] The determining unit 1403 is used to determine the traffic information of the point of interest to be identified in the current time period based on the trajectory information and correlation characteristics of the point of interest to be identified in the current time period.

[0433] In one example, the initial information also includes the access information of the point of interest to be identified in each time period within the historical time period, as well as the trajectory information of the point of interest to be identified in each time period within the historical time period; the point of interest to be identified has relevance features in each time period within the historical time period; wherein, the historical time period includes at least one time period.

[0434] In one example, unit 1403 is defined as including:

[0435] The acquisition module 14031 is used to acquire the feature image of the interest point to be identified; wherein, the feature image is used to indicate the features of the interest point.

[0436] The first processing module 14032 is used to process the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period to obtain the traffic information of the point of interest to be identified in the current time period.

[0437] In one example, module 14031 is retrieved, including:

[0438] The acquisition submodule 140311 is used to acquire the interest point information of the interest point to be identified; wherein the interest point information includes at least one feature information.

[0439] The first fusion submodule 140312 is used to perform data fusion based on the feature information in the interest point information of the interest point to be identified from the data source, so as to obtain the feature image of the interest point to be identified.

[0440] In one example, at least one feature includes road network information and trajectory data information; wherein, the road network information is the road network information of the location of the point of interest; and the trajectory data information is trajectory data information with the point of interest as the key location.

[0441] Module 14031 also includes:

[0442] The repair submodule 140313 is used to perform deviation repair processing on the trajectory data information based on road network information before the first fusion submodule 140312 performs data fusion based on the feature information in the interest point information of the interest point to be identified from the data source to obtain the feature image of the interest point to be identified, and obtains the repaired trajectory data information.

[0443] In one example, at least one feature includes one or more of the following: location information, map information, road network information, trajectory data information, and areal area information.

[0444] Among them, location information is the location information of the point of interest; map information is the map information of the location of the point of interest; road network information is the road network information of the location of the point of interest; trajectory data information is the trajectory data information with the point of interest as the key location; and areal region information is the areal region information of the point of interest.

[0445] In one example, the first processing module 14032 is specifically used to: process the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period based on the recognition model, so as to obtain the traffic information of the point of interest to be identified in the current time period.

[0446] In one example, determining unit 1403 also includes:

[0447] The second processing module 14033 is used to input the feature image of the point of interest to be identified into at least one convolutional neural network unit in the recognition model for processing before the first processing module 14032 processes the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period based on the recognition model to obtain the traffic information of the point of interest to be identified in the current time period, so as to obtain the processed feature image of the point of interest to be identified.

[0448] In one example, at least one convolutional neural network unit includes Q first convolutional neural network sub-units and MQ second convolutional neural network sub-units. The structure of the first convolutional neural network sub-units differs from the structure of the second convolutional neural network sub-units. Q is a positive integer greater than 1, M is a positive integer greater than 1, and M is greater than Q. The second processing module 14033 includes:

[0449] The first determining submodule 140331 is used to repeat the following steps until a first preset condition is met: inputting the feature image of the interest point to be identified into the qth first convolutional neural network subunit in the recognition model for processing to obtain the first feature map of the interest point to be identified; determining the first feature map of the interest point to be identified as the new feature image of the interest point to be identified, and determining that the value of q is q+1; wherein, the initial value of q is 1, and q is a positive integer greater than or equal to 1 and less than or equal to Q; the new feature image obtained when the first preset condition is met is the intermediate feature of the interest point to be identified.

[0450] The second determining submodule 140332 is used to repeat the following steps until the second preset condition is met: input the intermediate features of the interest point to be identified into the kth second convolutional neural network subunit in the recognition model for processing to obtain the second feature map of the interest point to be identified; determine the second feature map of the interest point to be identified as the new intermediate feature, and determine that the value of k is k+1; where the initial value of k is 1, and k is a positive integer greater than or equal to 1 and less than or equal to MQ.

[0451] Among them, the intermediate features obtained when the second preset condition is met are the processed feature images of the interest points to be identified.

[0452] In one example, the q-th first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer; the first determining submodule 140331, when inputting the feature image of the interest point to be identified into the q-th first convolutional neural network subunit in the recognition model for processing to obtain the first feature map of the interest point to be identified, is specifically used for:

[0453] The feature image of the interest point to be identified is input into the first convolutional layer for convolution processing to obtain the first convolutional feature; wherein, the first convolutional feature represents the convolutional feature of the feature image of the interest point to be identified; the first convolutional feature is input into the first activation layer for correction processing to obtain the corrected first convolutional feature; the corrected first convolutional feature is input into the pooling layer for feature compression processing to obtain the first feature map of the interest point to be identified; wherein, the first feature map of the interest point to be identified represents the feature image of the interest point to be identified.

[0454] In one example, the k-th second convolutional neural network subunit includes a second convolutional layer and a second activation layer; the second determining submodule 140332, when inputting the intermediate features of the interest point to be identified into the k-th second convolutional neural network subunit in the recognition model for processing to obtain the second feature map of the interest point to be identified, is specifically used for:

[0455] The intermediate features of the interest point to be identified are input into the second convolutional layer for convolution processing to obtain the second convolutional features; wherein, the second convolutional features represent the convolutional features of the intermediate features of the interest point to be identified; the second convolutional features are input into the second activation layer for correction processing to obtain the second feature map of the interest point to be identified.

[0456] In one example, the first processing module 14032 includes:

[0457] The second fusion submodule 140321 is used to input the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period into the first fully connected layer in the recognition model for feature fusion to obtain the fused features of the point of interest to be identified.

[0458] The prediction submodule 140322 is used to input the fused features of the interest point to be identified into the second fully connected layer in the recognition model for prediction processing, so as to obtain the traffic information of the interest point to be identified in the current time period.

[0459] In one example, the access information for the point of interest to be identified in each time period within the historical time period is determined based on each user's voice; wherein, the content of the user's voice includes the access information for the point of interest to be identified in each time period within the historical time period; the user's voice is obtained based on intelligent question answering.

[0460] In one example, the traffic information of the point of interest to be identified in each time period within the historical time period is the traffic information in each corrected speech; wherein, the corrected speech is the candidate speech with the highest value represented by the similarity information; the similarity information is the similarity between the user's speech and the candidate speech in the preset candidate speech set; the preset candidate speech set includes at least one candidate speech.

[0461] In one example, the similarity information is the similarity between the first audio information and the second audio information; wherein, the first audio information is the audio information of the user's speech, and the second audio information is the audio information of the candidate speech.

[0462] In one example, access information includes one or more of the following: accessible passageway, the orientation of the accessible passageway, the location of the point of interest, and the opening hours of the point of interest.

[0463] In one example, the apparatus provided in this embodiment further includes:

[0464] The navigation unit 1404 is used to generate a navigation path based on the traffic information of the point of interest to be identified in the current time period.

[0465] The control unit 1405 is used to control the autonomous vehicle to drive according to the navigation path.

[0466] The apparatus in this embodiment can execute the technical solutions in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.

[0467] Figure 15 This is a schematic diagram based on the ninth embodiment of the present disclosure, as shown below. Figure 15 As shown, the model training device 1500 for interest point information recognition provided in this embodiment includes:

[0468] The first acquisition unit 1501 is used to acquire a set of interest points to be trained; wherein, the set of interest points to be trained includes multiple interest points to be trained, and the interest points to be trained have access information and trajectory information. Access information is information that can indicate the accessibility of the interest point, and trajectory information represents the trajectory with the interest point as the key position.

[0469] The determining unit 1502 is used to determine the relevance features of the interest point to be trained based on the traffic information and trajectory information of the interest point to be trained; wherein, the relevance features characterize the correlation between the traffic information and trajectory information of the same interest point in the same time period.

[0470] Training unit 1503 is used to train the relevance features of the points of interest to be trained, so as to obtain the recognition model; wherein, the recognition model is used to identify the access information of the points of interest to be identified.

[0471] The apparatus in this embodiment can execute the technical solutions in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.

[0472] Figure 16 This is a schematic diagram based on the tenth embodiment of the present disclosure, as shown below. Figure 16 As shown, the model training device 1600 for interest point information recognition provided in this embodiment includes:

[0473] The first acquisition unit 1601 is used to acquire a set of interest points to be trained; wherein, the set of interest points to be trained includes multiple interest points to be trained, and the interest points to be trained have access information and trajectory information. Access information is information that can indicate the accessibility of the interest point, and trajectory information represents the trajectory with the interest point as the key position.

[0474] The determining unit 1602 is used to determine the relevance features of the interest point to be trained based on the traffic information and trajectory information of the interest point to be trained; wherein, the relevance features characterize the correlation between the traffic information and trajectory information of the same interest point in the same time period.

[0475] Training unit 1603 is used to train the relevance features of the points of interest to be trained, so as to obtain the recognition model; wherein, the recognition model is used to identify the access information of the points of interest to be identified.

[0476] In one example, the point of interest to be trained has traffic information and trajectory information for each time period within a preset time period set; the point of interest to be trained has correlation features for each time period within the preset time period set; wherein, the preset time period set includes at least one time period.

[0477] In one example, the point of interest to be trained has real traffic information, which is the traffic information of the point of interest to be trained in the last time period within a preset time period set; training unit 1603 includes:

[0478] The processing subunit 16031 is used to input the relevance features of the interest point to be trained into the initial model to obtain the predicted traffic information of the interest point to be trained; wherein, the predicted traffic information is the traffic information of the interest point to be trained in the last time period.

[0479] The adjustment subunit 16032 is used to adjust the parameters of the initial model based on the predicted traffic information of the interest point to be trained and the actual traffic information of the interest point to be trained, so as to obtain the recognition model.

[0480] In one example, processing subunit 16031 includes:

[0481] The acquisition module 160311 is used to acquire the feature image of the interest point to be trained; wherein, the feature image is used to indicate the features of the interest point.

[0482] The first processing module 160312 is used to input the feature image of the interest point to be trained and the correlation features of the interest point to be trained into the initial model to obtain the prediction information of the interest point to be trained.

[0483] In one example, retrieving module 160311 includes:

[0484] The acquisition submodule 1603111 is used to acquire the interest point information of the interest points to be trained; wherein, the interest point information includes at least one feature information;

[0485] The first fusion submodule 1603112 is used to perform data fusion based on the feature information in the interest point information of the interest point to be trained from the data source, so as to obtain the feature image of the interest point to be trained.

[0486] In one example, at least one feature includes road network information and trajectory data information; wherein, the road network information is the road network information of the location of the point of interest; and the trajectory data information is trajectory data information with the point of interest as the key location.

[0487] Module 160311 also includes:

[0488] The repair submodule 1603113 is used to perform deviation repair processing on the trajectory data information based on road network information before the first fusion submodule 1603112 performs data fusion based on the feature information in the interest point information of the interest point to be trained based on the data source to obtain the feature image of the interest point to be trained, and obtains the repaired trajectory data information.

[0489] In one example, at least one feature includes one or more of the following: location information, map information, road network information, trajectory data information, and areal area information.

[0490] Among them, location information is the location information of the point of interest; map information is the map information of the location of the point of interest; road network information is the road network information of the location of the point of interest; trajectory data information is the trajectory data information with the point of interest as the key location; and areal region information is the areal region information of the point of interest.

[0491] In one example, processing subunit 16031 also includes:

[0492] The second processing module 160313 is used to input the feature image of the interest point to be trained and the correlation features of the interest point to be trained into the initial model before the first processing module 160312 inputs them into the initial model to obtain the predicted traffic information of the interest point to be trained, and then inputs them into the initial model to obtain the processed feature image of the interest point to be trained.

[0493] In one example, at least one convolutional neural network unit includes n first convolutional neural network sub-units and mn second convolutional neural network sub-units. The structure of the first convolutional neural network sub-units differs from the structure of the second convolutional neural network sub-units. n is a positive integer greater than 1, m is a positive integer greater than 1, and m is greater than n. The second processing module 160313 includes:

[0494] The first determining submodule 1603131 is used to repeat the following steps until the first preset condition is met: inputting the feature image of the interest point to be trained into the i-th first convolutional neural network subunit in the initial model for processing to obtain the first feature map of the interest point to be trained; determining the first feature map of the interest point to be trained as the new feature image of the interest point to be trained, and determining that the value of i is i+1; where the initial value of i is 1, and i is a positive integer greater than or equal to 1 and less than or equal to n; the new feature image obtained when the first preset condition is met is the intermediate feature of the interest point to be trained.

[0495] The second determining submodule 1603132 is used to repeat the following steps until the second preset condition is met: input the intermediate features of the interest point to be trained into the j-th second convolutional neural network subunit in the initial model for processing to obtain the second feature map of the interest point to be trained; determine the second feature map of the interest point to be trained as the new intermediate feature, and determine that the value of j is j+1; where the initial value of j is 1, and j is a positive integer greater than or equal to 1 and less than or equal to mn.

[0496] Among them, the intermediate features obtained when the second preset condition is met are the processed feature images of the interest points to be trained.

[0497] In one example, the i-th first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer; the first determining submodule 1603131, when inputting the feature image of the interest point to be trained into the i-th first convolutional neural network subunit in the initial model for processing to obtain the first feature map of the interest point to be trained, is specifically used for:

[0498] The feature image of the interest point to be trained is input into the first convolutional layer for convolution processing to obtain the first convolutional feature; wherein, the first convolutional feature represents the convolutional feature of the feature image of the interest point to be trained; the first convolutional feature is input into the first activation layer for correction processing to obtain the corrected first convolutional feature; the corrected first convolutional feature is input into the pooling layer for feature compression processing to obtain the first feature map of the interest point to be trained; wherein, the first feature map of the interest point to be trained represents the feature image of the interest point to be trained.

[0499] In one example, the j-th second convolutional neural network subunit includes a second convolutional layer and a second activation layer; the second determining submodule 1603132, when inputting the intermediate features of the interest point to be trained into the j-th second convolutional neural network subunit in the initial model for processing to obtain the second feature map of the interest point to be trained, is specifically used for:

[0500] The intermediate features of the interest points to be trained are input into the second convolutional layer for convolution processing to obtain the second convolutional features; wherein, the second convolutional features represent the convolutional features of the intermediate features of the interest points to be trained; the second convolutional features are input into the second activation layer for correction processing to obtain the second feature map of the interest points to be trained.

[0501] In one example, the first processing module 160312 includes:

[0502] The second fusion submodule 1603121 is used to input the feature image of the interest point to be trained and the correlation features of the interest point to be trained into the first fully connected layer in the initial model for feature fusion to obtain the fused features of the interest point to be trained.

[0503] The prediction submodule 1603122 is used to input the fused features of the interest points to be trained into the second fully connected layer in the initial model for prediction processing, so as to obtain the prediction information of the interest points to be trained.

[0504] In one example, the first acquisition unit 1601 includes:

[0505] The first acquisition subunit 16011 is used to acquire user voice based on intelligent question answering; wherein, the content of the user voice includes access information of the interest points to be trained;

[0506] The subunit 16012 is used to determine the passage information corresponding to the user's voice based on the user's voice.

[0507] The second acquisition subunit 16013 is used to acquire the trajectory information of the interest points to be trained corresponding to the user's voice, so as to obtain the interest points to be trained in the set of interest points to be trained.

[0508] In one example, the first acquisition subunit 16011 includes:

[0509] The generation module 160111 is used to obtain user intent and generate question speech based on user intent and preset question text corresponding to the interest points to be trained.

[0510] The sending module 160112 is used to send voice messages with questions to the user.

[0511] The receiving module 160113 is used to receive user voice feedback.

[0512] In one example, subunit 16012 is identified as including:

[0513] The first determining module 160121 is used to determine the similarity between the user's voice and candidate voices in a preset candidate voice set, and to obtain similarity information of the candidate voices, wherein the preset candidate voice set includes at least one candidate voice.

[0514] The second determining module 160122 is used to determine the candidate speech with the highest similarity value represented by the similarity information, for correcting the speech.

[0515] The third determining module 160123 is used to determine the passage information in the corrected speech, which is the passage information corresponding to the user's speech.

[0516] In one example, the first determining module 160121 is specifically used for:

[0517] Extract the first audio information of the user's speech and extract the second audio information of the candidate speech; determine the similarity between the first audio information and the second audio information to obtain the similarity information of the candidate speech.

[0518] In one example, access information includes one or more of the following: accessible passageway, the orientation of the accessible passageway, the location of the point of interest, and the opening hours of the point of interest.

[0519] The apparatus in this embodiment can execute the technical solutions in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.

[0520] Figure 17 This is a schematic diagram based on the eleventh embodiment of the present disclosure, as shown below. Figure 17 As shown, the electronic device 1700 in this embodiment may include a processor 1701 and a memory 1702.

[0521] Memory 1702 is used to store programs. Memory 1702 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; memory may also include non-volatile memory, such as flash memory. Memory 1702 is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc. The computer programs, computer instructions, etc., can be partitioned and stored in one or more memories 1702. Furthermore, the computer programs, computer instructions, data, etc., can be accessed by processor 1701.

[0522] The aforementioned computer programs and instructions can be stored in one or more partitions of memory 1702. Furthermore, the aforementioned computer programs and instructions can be invoked by processor 1701.

[0523] Processor 1701 is configured to execute computer programs stored in memory 1702 to implement the various steps in the methods described in the above embodiments.

[0524] For details, please refer to the relevant descriptions in the preceding method embodiments.

[0525] The processor 1701 and memory 1702 can be independent structures or integrated structures. When the processor 1701 and memory 1702 are independent structures, the memory 1702 and the processor 1701 can be coupled together via bus 1703.

[0526] The electronic device in this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principle are the same, and will not be repeated here.

[0527] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0528] According to embodiments of this disclosure, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the scheme provided in any of the above embodiments.

[0529] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.

[0530] Figure 18 A schematic block diagram of an example electronic device 1800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0531] like Figure 18 As shown, device 1800 includes a computing unit 1801, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1802 or a computer program loaded from storage unit 1808 into random access memory (RAM) 1803. The RAM 1803 may also store various programs and data required for the operation of device 1800. The computing unit 1801, ROM 1802, and RAM 1803 are interconnected via bus 1804. Input / output (I / O) interface 1805 is also connected to bus 1804.

[0532] Multiple components in device 1800 are connected to I / O interface 1805, including: input unit 1806, such as keyboard, mouse, etc.; output unit 1807, such as various types of monitors, speakers, etc.; storage unit 1808, such as disk, optical disk, etc.; and communication unit 1809, such as network card, modem, wireless transceiver, etc. Communication unit 1809 allows device 1800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0533] The computing unit 1801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1801 performs the various methods and processes described above, such as methods for identifying points of interest or methods for training models applied to the identification of points of interest. For example, in some embodiments, the methods for identifying points of interest or methods for training models applied to the identification of points of interest may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1800 via ROM 1802 and / or communication unit 1809. When the computer program is loaded into RAM 1803 and executed by the computing unit 1801, one or more steps of the methods for identifying points of interest or methods for training models applied to the identification of points of interest described above may be performed. Alternatively, in other embodiments, the computing unit 1801 may be configured by any other suitable means (e.g., by means of firmware) to perform an information recognition method for points of interest or a model training method applied to the information recognition of points of interest.

[0534] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0535] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0536] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0537] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0538] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0539] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0540] According to embodiments of this disclosure, this disclosure also provides an autonomous driving vehicle, in which the electronic devices provided in the above embodiments are provided.

[0541] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0542] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for identifying points of interest, the method comprising: Obtain initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period and the access information of the point of interest to be identified in each time period in the historical time period, wherein, the trajectory information represents the trajectory with the point of interest as the key location, and the access information includes one or more of the following: accessible passage entrances, the orientation of accessible passage entrances, the location of the point of interest, and the opening time of the point of interest; The access information of the point of interest to be identified in each time period within the historical time period is determined based on each user's voice. The content of the user's voice includes the access information of the point of interest to be identified in each time period within the historical time period. The user's voice is obtained based on an intelligent question-and-answer method. The intelligent question-answering method includes: obtaining user intent; determining preset question text corresponding to both the point of interest to be identified and the user intent based on the traffic information to be obtained in the current time period for the point of interest to be identified; generating question voice based on the user intent and the preset question text, wherein each preset question text includes text information asking a question about the point of interest to be identified; sending the question voice to the user and receiving user voice feedback. The traffic information of the point of interest to be identified in each time period within the historical time period is the traffic information in each corrected speech; the corrected speech is the candidate speech with the highest value represented by the similarity information; the similarity information is the similarity between the user's speech and the candidate speech in the preset candidate speech set; the preset candidate speech set includes at least one candidate speech, and each candidate speech set corresponds to each geographical region. The relevance features of the point of interest to be identified are obtained; wherein, the relevance features are obtained by analyzing the correlation between the access information and trajectory information of the point of interest to be identified in the same time period within a historical time period, and are used to characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest. Based on the trajectory information of the point of interest to be identified in the current time period and the correlation features, the traffic information of the point of interest to be identified in the current time period is determined.

2. The method according to claim 1, wherein, The initial information also includes trajectory information of the point of interest to be identified in each time period within the historical time period; the point of interest to be identified has correlation characteristics in each time period within the historical time period; The historical time period includes at least one time period.

3. The method according to claim 2, wherein, Based on the trajectory information of the point of interest to be identified in the current time period and the correlation features, the traffic information of the point of interest to be identified in the current time period is determined, including: Obtain the feature image of the point of interest to be identified; wherein the feature image is used to indicate the features of the point of interest; The feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period are processed to obtain the traffic information of the point of interest to be identified in the current time period.

4. The method according to claim 3, wherein, Obtaining the feature image of the interest point to be identified includes: Obtain the interest point information of the interest point to be identified; wherein the interest point information includes at least one feature information; Based on the feature information in the interest point information of the interest point to be identified from the data source, data fusion is performed to obtain the feature image of the interest point to be identified.

5. The method according to claim 4, wherein, The at least one feature information includes road network information and trajectory data information; wherein, the road network information is the road network information of the location of the point of interest; and the trajectory data information is trajectory data information with the point of interest as the key location. Before performing data fusion on the feature information of the interest point to be identified based on the data source to obtain the feature image of the interest point to be identified, the method further includes: Based on the road network information, deviation correction processing is performed on the trajectory data information to obtain the corrected trajectory data information.

6. The method according to claim 4, wherein, The at least one feature information includes one or more of the following: location information, map information, road network information, trajectory data information, and areal region information; Wherein, the location information is the location information of the point of interest; the map information is the map information of the location of the point of interest; the road network information is the road network information of the location of the point of interest; the trajectory data information is the trajectory data information with the point of interest as the key location; and the areal region information is the areal region information of the point of interest.

7. The method according to any one of claims 3-6, wherein, The feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period are processed to obtain the traffic information of the point of interest to be identified in the current time period, including: Based on the recognition model, the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period are processed to obtain the traffic information of the point of interest to be identified in the current time period.

8. The method according to claim 7, further comprising, before processing the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period based on the recognition model to obtain the traffic information of the point of interest to be identified in the current time period: The feature image of the interest point to be identified is input into at least one convolutional neural network unit in the recognition model for processing to obtain the processed feature image of the interest point to be identified.

9. The method according to claim 8, wherein, The at least one convolutional neural network unit includes Q first convolutional neural network sub-units and MQ second convolutional neural network sub-units. The structure of the first convolutional neural network sub-units differs from the structure of the second convolutional neural network sub-units. Q is a positive integer greater than 1, M is a positive integer greater than 1, and M is greater than Q. The feature image of the interest point to be identified is input into at least one convolutional neural network unit in the recognition model for processing to obtain the processed feature image of the interest point to be identified, including: Repeat the following steps until the first preset condition is met: input the feature image of the interest point to be identified into the q-th first convolutional neural network subunit in the recognition model for processing to obtain the first feature map of the interest point to be identified; determine the first feature map of the interest point to be identified as the new feature image of the interest point to be identified, and determine that the value of q is q+1; wherein, the initial value of q is 1, and q is a positive integer greater than or equal to 1 and less than or equal to Q; the new feature image obtained when the first preset condition is met is the intermediate feature of the interest point to be identified; Repeat the following steps until the second preset condition is met: input the intermediate features of the interest point to be identified into the kth second convolutional neural network subunit in the recognition model for processing to obtain the second feature map of the interest point to be identified; determine that the second feature map of the interest point to be identified is a new intermediate feature, and determine that the value of k is k+1; where the initial value of k is 1, and k is a positive integer greater than or equal to 1 and less than or equal to MQ; The intermediate features obtained when the second preset condition is met are the processed feature images of the interest points to be identified.

10. The method according to claim 9, wherein, The q-th first convolutional neural network subunit includes a first convolutional layer, a first activation layer, and a pooling layer; the feature image of the interest point to be identified is input into the q-th first convolutional neural network subunit in the recognition model for processing to obtain the first feature map of the interest point to be identified, including: The feature image of the interest point to be identified is input into the first convolutional layer for convolution processing to obtain the first convolutional feature; wherein, the first convolutional feature represents the convolutional feature of the feature image of the interest point to be identified; The first convolutional feature is input into the first activation layer for correction processing to obtain the corrected first convolutional feature; The corrected first convolutional features are input into the pooling layer for feature compression processing, resulting in the first feature map of the interest point to be identified; wherein, the first feature map of the interest point to be identified represents the feature image of the interest point to be identified.

11. The method according to claim 9 or 10, wherein, The k-th second convolutional neural network subunit includes a second convolutional layer and a second activation layer; the intermediate features of the interest point to be identified are input into the k-th second convolutional neural network subunit in the recognition model for processing to obtain the second feature map of the interest point to be identified, including: The intermediate features of the interest point to be identified are input into the second convolutional layer for convolution processing to obtain the second convolutional features; wherein, the second convolutional features represent the convolutional features of the intermediate features of the interest point to be identified; The second convolutional features are input into the second activation layer for correction processing to obtain the second feature map of the interest point to be identified.

12. The method according to claim 11, wherein, Based on the recognition model, the feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period are processed to obtain the traffic information of the point of interest to be identified in the current time period, including: The feature image of the point of interest to be identified, the correlation features of the point of interest to be identified, and the trajectory information of the point of interest to be identified in the current time period are input into the first fully connected layer in the identification model for feature fusion to obtain the fused features of the point of interest to be identified. The fused features of the interest point to be identified are input into the second fully connected layer of the identification model for prediction processing to obtain the traffic information of the interest point to be identified in the current time period.

13. The method according to claim 1, wherein, The similarity information is the similarity between the first audio information and the second audio information; Wherein, the first audio information is the audio information of the user's voice, and the second audio information is the audio information of the candidate voice.

14. The method of claim 13, further comprising: Based on the traffic information of the points of interest to be identified in the current time period, a navigation path is generated; The autonomous vehicle is controlled to drive according to the navigation path.

15. A model training method for information recognition of points of interest, the method comprising: Obtain a set of points of interest to be trained; wherein the set of points of interest to be trained includes multiple points of interest to be trained, and each point of interest to be trained has access information and trajectory information for each time period within a preset time period set. The access information is information that can indicate the accessibility of the point of interest, and the trajectory information represents the trajectory with the point of interest as the key location. The access information includes one or more of the following: accessible passage entrances, the orientation of accessible passage entrances, the location of the point of interest, and the opening time of the point of interest. By analyzing the correlation between the traffic information and trajectory information of the point of interest to be trained in the same historical time period, the correlation characteristics of the point of interest to be trained are determined; wherein, the correlation characteristics represent the correlation between the traffic information and trajectory information of the same point of interest in the same time period. The relevance features of the points of interest to be trained are processed to obtain a recognition model; wherein, the recognition model is used to identify the access information of the points of interest to be identified; The access information is the access information corresponding to the user's voice. Determine the access information corresponding to the user's voice, including: Determine the similarity between the user's voice and candidate voices in a preset candidate voice set to obtain similarity information of the candidate voices, wherein the preset candidate voice set includes at least one candidate voice, and each candidate voice set corresponds to a geographical region; The candidate speech with the highest similarity value is identified as the corrected speech. The access information in the corrected speech is determined to be the access information corresponding to the user's speech; Obtain the set of interest points to be trained, including: User voice is acquired using an intelligent question-and-answer method; wherein, the content of the user voice includes access information for the points of interest to be trained. Based on the user's voice, determine the passage information corresponding to the user's voice; and obtain the trajectory information of the interest points to be trained corresponding to the user's voice, so as to obtain the interest points to be trained in the set of interest points to be trained. User voice is obtained based on intelligent question answering, including: Obtain user intent; For the traffic information to be obtained for the current time period of the point of interest to be identified, a preset question text corresponding to both the point of interest to be identified and the user intent is determined; Based on the user intent and the preset question text, generate question speech, whereby each preset question text includes text information about a question asked about a point of interest to be identified. Send the question to the user via voice message and receive user feedback via voice message.

16. The method according to claim 15, wherein, The points of interest to be trained also have trajectory information for each time period within the preset time period set; the points of interest to be trained have correlation characteristics for each time period within the preset time period set. The preset time period set includes at least one time period.

17. The method according to claim 16, wherein, The point of interest to be trained has real traffic information, which is the traffic information of the point of interest to be trained in the last time period within the preset time period set; The relevance features of the points of interest to be trained are processed to obtain a recognition model, including: The relevance features of the interest point to be trained are input into the initial model to obtain the predicted traffic information of the interest point to be trained; wherein, the predicted traffic information is the traffic information of the interest point to be trained in the last time period; Based on the predicted traffic information of the interest points to be trained and the actual traffic information of the interest points to be trained, the parameters of the initial model are adjusted to obtain the recognition model.

18. The method according to claim 17, wherein, The relevance features of the interest points to be trained are input into the initial model to obtain the traffic information for the prediction of the interest points to be trained, including: Obtain the feature image of the interest point to be trained; wherein the feature image is used to indicate the features of the interest point; The feature image of the interest point to be trained and its relevance features are input into the initial model to obtain the predicted access information of the interest point to be trained.

19. The method according to claim 15, wherein, Determine the similarity between the user's speech and candidate speech in a preset candidate speech set to obtain the similarity information of the candidate speech, including: Extract the first audio information of the user's speech, and extract the second audio information of the candidate speech; The similarity between the first audio information and the second audio information is determined to obtain the similarity information of the candidate speech.

20. An information identification device for points of interest, the device comprising: The first acquisition unit is used to acquire initial information of the point of interest to be identified; wherein, the initial information includes the trajectory information of the point of interest to be identified in the current time period and the access information of the point of interest to be identified in each time period in the historical time period, wherein, the trajectory information represents the trajectory with the point of interest as the key location, and the access information includes one or more of the following: accessible passage entrances, the orientation of the accessible passage entrances, the location of the point of interest, and the opening time of the point of interest; the access information of the point of interest to be identified in each time period in the historical time period is determined based on each user's voice, and the content of the user's voice includes the access information of the point of interest to be identified in each time period in the historical time period; the user's voice is acquired based on an intelligent question-and-answer method; The intelligent question-answering method includes: obtaining user intent; determining preset question text corresponding to both the point of interest to be identified and the user intent based on the traffic information to be obtained in the current time period for the point of interest to be identified; generating question voice based on the user intent and the preset question text, wherein each preset question text includes text information asking a question about the point of interest to be identified; sending the question voice to the user and receiving user voice feedback. The traffic information of the point of interest to be identified in each time period within the historical time period is the traffic information in each corrected speech; the corrected speech is the candidate speech with the highest value represented by the similarity information; the similarity information is the similarity between the user's speech and the candidate speech in the preset candidate speech set; the preset candidate speech set includes at least one candidate speech, and each candidate speech set corresponds to each geographical region. The second acquisition unit is used to acquire the relevance features of the point of interest to be identified; wherein, the relevance features are obtained by analyzing the correlation between the access information and trajectory information of the point of interest to be identified in the same time period within a historical time period, and are used to characterize the correlation between the access information and trajectory information of the same point of interest in the same time period; the access information is information that can indicate the accessibility of the point of interest. The determining unit is used to determine the traffic information of the point of interest to be identified in the current time period based on the trajectory information of the point of interest to be identified in the current time period and the correlation features.

21. A model training device for information recognition of points of interest, the device comprising: The first acquisition unit is used to acquire a set of interest points to be trained; wherein, the set of interest points to be trained includes multiple interest points to be trained, and each interest point to be trained has access information and trajectory information for each time period within a preset time period set. The access information is information that can indicate the accessibility of the interest point, and the trajectory information represents the trajectory with the interest point as the key location. The access information includes one or more of the following: accessible passage entrance, orientation of the accessible passage entrance, location of the interest point, and opening time of the interest point. The determining unit is used to determine the relevance features of the point of interest to be trained by analyzing the correlation between the traffic information and trajectory information of the same point of interest in the same historical time period; wherein, the relevance features characterize the correlation between the traffic information and trajectory information of the same point of interest in the same time period. The training unit is used to train the relevance features of the interest points to be trained to obtain a recognition model; wherein, the recognition model is used to identify the access information of the interest points to be identified. The access information is the access information corresponding to the user's voice. The first determining module is used to determine the similarity between the user's voice and candidate voices in a preset candidate voice set, and to obtain the similarity information of the candidate voices. The preset candidate voice set includes at least one candidate voice, and each candidate voice set corresponds to a geographical region. The second determining module is used to determine the candidate speech with the highest similarity value represented by the similarity information, which is the corrected speech. The third determining module is used to determine the access information in the corrected speech, which is the access information corresponding to the user's speech; The first acquisition unit includes: The first acquisition subunit is used to acquire user voice based on intelligent question answering; wherein, the content of the user voice includes access information of the interest points to be trained; A determining subunit is used to determine the passage information corresponding to the user's voice based on the user's voice. The second acquisition subunit is used to acquire the trajectory information of the interest points to be trained corresponding to the user's voice, so as to obtain the interest points to be trained in the set of interest points to be trained. The first acquisition subunit includes: The generation module is used to obtain user intent; for the traffic information to be obtained for the current time period of the point of interest to be identified, determine a preset question text that corresponds to both the point of interest to be identified and the user intent; and generate question speech based on the user intent and the preset question text, wherein each preset question text includes text information of a question about the point of interest to be identified. The sending module is used to send the question to the user via voice. The receiving module is used to receive user voice feedback.

22. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-19.

23. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-19.

24. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-19.

25. An autonomous vehicle, wherein the autonomous vehicle is equipped with the electronic equipment as described in claim 22.

Citation Information

Patent Citations

  • Method, device and equipment for determining passing state of gate and computer storage medium

    CN111858800A

  • Map guide point determination method and device, equipment, storage medium and program product

    CN114329246A