Blog posts

Digging into the data: The Cropster Data Project

In part 2 of our series on the Cropster Data Project we took a closer look at how we process, protect and secure the data we used for the project. In this third part, we'll take a look at the results of our research and the implications it has for the coffee community.

Machine and Human Learning

Before digging into the details, we had to convert the data into a form that could be easily analyzed. During this phase, our data engineering team worked closely with the researchers to mold the data into the right form, eliminate statistical outliers and decide which data subsets were the most relevant to the research questions. During this phase the data was also cleaned and pre-processed.

After the data cleaning, the research team tried to identify which roasting parameters have the largest influence on the quality of the roasted coffee (as measured by cupping score). In this case we used temperature time series data and the corresponding metadata and tried to find relationships with the cupping score. In total we analyzed more than 40 features such as mean and average temperatures in certain phases, deviations, or the skewness and kurtosis of the roast curves. As more than 40 features are hard to visualize and interpret at once, we reduced the scope of our data analysis with a principal component analysis. As we expected, due to the many relevant variables of roasting, we could not find any significant patterns with this first approach, though it did point us in the right direction going forward.

We then analyzed the similarity of a roast curve with its corresponding profile curve. Although this sounds relatively easy, it is actually quite complex. Since roast curves are never exactly the same length, we had to use shape matching algorithms to align the roast curve with its profile curve without losing too much information. Unfortunately, the difference between the roast and profile curve did not show a direct correlation with the cupping score. 

Next, we simplified our approach by defining that "good roasts" have the highest possible normalized score while "bad roasts" have any score that is smaller than the highest possible score. We applied state-of-the-art machine learning algorithms to the simplified data set and were then able to detect "good or bad roasts" with a statistically significant level of accuracy.

Then, we intensified our analysis on the detection of flavors that a roast was going to have. The university team developed a new representation of the problem space which allowed them to train machine learning algorithms to predict the flavor labels based on the temperature curve. This task required intensive preprocessing of the labels in collaboration with one of our roasting experts. The results did not only allow the prediction of flavor labels of a roast based on its curve to a significant degree, but also showed for instance that bitter tastes form during roasts that remain longer in a medium temperature range (ca. 125°C - 175 °C) before they reach high ranges (>220C°) later during the roast.

The results of the first project were refined in a second project. In the second project, we improved the mapping between flavors and the temperature curve by adding sentiment labels to the flavor labels and by combining the labels in a systematic manner. This allowed us to reduce the potential classes and eliminate uncommon labels from the target dataset. With the second dataset we were able to expand the prototype for predicting multiple combinations of labels at once.

Outlook

The Cropster Data Project has already shown us that there is much more to discover. Even with a few small scale research projects, we have already found and visualized relationships between roasting data and quality of the final product. Some of the ideas worked well, and though some proved not to be worth pursuing, we are just getting started and plan to continue our collaboration with enthusiastic researchers from different domains. Many aspects of coffee roasting remain mysterious and leave us with more questions than answers due to the new knowledge we gained. These questions invite further analysis of the coffee roasting process and we are excited to continue that research!

 

更多最新消息

New releases

最新發布:鍋間流程 - 穩定烘焙的秘密

我們很興奮地向大家宣布,Cropster 的又一新功能 ——「鍋間流程」來了!減輕烘焙師們的工作量,使最關鍵的業務信息愈加易於收集、訪問及分析。

浏览更多
Roastery   -   Blog posts

您的最佳升温速度 (RoR) 是什么?

在 Cropster,我们能收到很多关于升温速度的问题。 作为热爱咖啡的技术人员,我们对此有诸多了解。 我们还知道,我们在技术方面的专业知识只能帮烘焙师们这么多。 他们还需要更深入的咖啡专业知识。 因此,我们找到世界领先的咖啡烘焙专家 Anne Cooper、Rob Hoos 和 Scott Rao 来帮助我们解答关于升温速度的核心问题,了解烘焙师如何运用 Cropster…

浏览更多
New releases

烘焙与人工智能相结合的下一个重要功能来了——一爆预测!

Cropster 对 Roasting Intelligence 智能烘焙 4 进行最新更新,在人工智能 (AI) 的基础上推出一爆预测功能,该功能指示一爆的可能发生时间,方便您据以进行烘焙规划。 继续阅读全文。

浏览更多

订阅我们的新闻邮件

了解为以下人士提供的解决方案

Here should be a form, apparently your browser blocks our forms.

Do you use an adblocker? If so, please try turning it off and reload this page.