Data Blog by Lizeo
First and foremost, to efficiently track the prices charged on your market, you must first define a scope for brands and data sources. For example, this includes e-shops in the case of web data collection. Next, you need to determine, in detail, which competitors’ products you want to track as a priority. At this stage, it is important to ensure that the available data sources are sufficiently homogeneous and granular. Consequently, this guarantees a precise comparison with your references. Finally, make sure that the names and technical descriptions of competitors’ products are unified between your different data sources. This process is known as product matching. Ultimately, it is at this last step that your product reference system comes into play. Naturally, it must be enhanced and validated by business experts.
Fundamentally, in order to decipher the competitive data collected, you must use a clean product reference system. Specifically, to obtain the most representative vision of the market possible, you need to collect and unify the data. This applies to different sources, whether they are online or offline. However, there is a problem: the nomenclature, description, and characteristics of the same product generally vary from one source to another. As a result, this makes the analysis imprecise, or even impossible due to duplicate data or holes in the data.
To effectively address this difficulty caused by the diversity of collection points, a data matching phase must be integrated. This must happen upstream of any analysis. In practice, this step will link the price collected from a reference found on a source (site A or price list A) with the price of this same reference on another source (site B or price list B). Moreover, the reference must also be identified in the certified product database. This guarantees the relevance of the prices that are now attached to it.
Overall, the operation is simple. Each piece of product information collected during the collection of prices must be sorted automatically into three categories:
First, Unknown: The data can refer to a product that does not exist or to a new product to be created from a reliable source. In this case, we will search for official documents certifying the existence or not of the product. This allows us to complete the reference system and retroactively match the product.
Second, Known: The pricing data must be associated with the corresponding product.
Third, False: The product must be blacklisted. For example, this is the case when a piece of technical information concerning a product available in a source is trivially false.
By automating this sorting operation, it cleans the data by excluding incorrect information and duplicates. Ultimately, this provides real time savings. As a result, the “matched” data is stored in a clean database by associating a single line of information per product.
Are your products complex? Do they have many technical features? If so, building and using a reference database will simplify your matching operations.
To guarantee the quality and completeness of this product database, a continuous enhancement process is required. Importantly, this must be supervised by business experts.
Furthermore, each element of the reference must be traced and created from official sources only. These include sales catalogs and manufacturers’ websites. Additionally, the detailed qualification of each product assisted by machine learning should be validated by experts in the field.
However, note that to carry out matching, efficient reconciliation, or linking, you will need an overall view of the product offers available on your market. Consequently, the level of precision of this catalog will determine the precision and depth of the price analysis carried out afterwards. Now, it only remains to define and select the discriminating criteria on your market. Thanks to your teams, this makes your comparative analysis as relevant as possible.
Traceability: Each product or attribute creation must come from an official source which will be archived.
Representativity: The level of qualification and clustering of competitors’ products must be defined by business experts. Moreover, it must reflect an unbiased vision of the market.
Completeness: The more complete the reference, the richer the analysis. Therefore, this involves a balance between automatic qualification and manual enhancement.
cleaning and matching operations are essential.