1How do you explain foundational metrics like conversion rates or statistical significance to non-technical business stakeholders without using technical jargon?
When communicating foundational metrics to non-technical stakeholders, the key is to translate abstract formulas and statistical mechanics into intuitive user stories, decision confidence, and business risk. For conversion rate, rather than presenting a bare ratio or abstract percentage, frame it around concrete user counts: 'Out of every 100 people who landed on the checkout page, 5 completed a purchase.' This grounds the metric directly in observable customer behavior. For statistical significance, avoid formal null hypothesis terminology like alpha levels or rejection regions. Instead, explain it as confidence that an observed uplift is real rather than random noise or lucky timing. For instance: 'A statistically significant result means we are 95% confident this improvement reflects a genuine change in user behavior rather than random fluke. Rolling this out gives us high confidence of a positive outcome instead of reacting to random noise.'
## Experiment Results: New Checkout Flow
- Baseline Conversion: 5.0% (5 out of 100 visitors buy)
- Variant Conversion: 5.8% (5.8 out of 100 visitors buy, a +16% relative lift)
- Decision Confidence: 96% confidence (p = 0.04)
- Business Takeaway: The risk that this lift was just random chance is under 4%. Rolling this out is estimated to bring +$45k in monthly revenue.
-- Inner join silently shrinks denominator to purchasers only
SELECT COUNT(o.order_id) * 1.0 / COUNT(u.user_id) AS inner_conv_rate
FROM users u
INNER JOIN orders o ON u.user_id = o.user_id;
-- Left join preserves full user base for accurate metric computation
SELECT COUNT(o.order_id) * 1.0 / COUNT(u.user_id) AS true_conv_rate
FROM users u
LEFT JOIN orders o ON u.user_id = o.user_id;
3In product experimentation, what does randomization achieve that simple before-after comparison usually cannot, and how would you explain the core causal claim an A/B test is designed to support?
In product experimentation, a simple before-after comparison compares metrics across different time periods, which confounds the effect of a feature change with external temporal factors such as seasonality, day-of-week patterns, concurrent marketing campaigns, macro trends, and natural user maturation. Randomization assigns eligible units simultaneously to treatment and control groups, balancing both observed and unobserved confounding variables across groups in expectation. This establishes internal validity. The core causal claim supported by an A/B test relies on counterfactual reasoning: because the control group is subjected to the exact same external conditions over the exact same time window, it serves as an empirical estimate of the counterfactual—what would have happened to the treatment group had they not received the feature. Therefore, any statistically significant difference in outcomes can be causally attributed to the treatment intervention.
# Naive Before-After Comparison:
# Effect_estimate = Metric_t1 (Holiday Launch) - Metric_t0 (Pre-Holiday Baseline)
# Problem: Lift is confounded by holiday shopping surge and marketing spend.
# Randomized A/B Test:
# Effect_estimate = E[Metric_Holiday | Treatment] - E[Metric_Holiday | Control]
# Solution: Both arms experience the exact same external shocks simultaneously.
4正式なモデリングや実験を行う前に、EDA (Exploratory Data Analysis) は通常何を達成することを目的としていますか?また、EDAと確証的データ分析(confirmatory analysis)はどのように区別されますか?
探索的データ分析(EDA: Exploratory Data Analysis)は、正式な統計モデリングや実験の前に、データセットの潜在的な構造を把握し、パターンを発見し、異常値やデータ品質の問題を特定し、分布の前提条件を評価し、仮説を生成することを目的とした自由度の高いプロセスです。
EDAと確証的データ分析の根本的な違いは、その目的と方法論にあります。EDAは「仮説を生成する」ための柔軟な探索的アプローチであり、記述統計の要約値、相関、視覚化を用いて、厳密な前提条件に縛られることなくデータが示唆する内容を探ります。これに対して確証的データ分析(仮説検定やA/Bテストの評価など)は「仮説を検証する」ための構造化された推測統計的アプローチであり、統計的過誤率(第1種の過誤など)を制御しながら、事前に設定された反証可能な仮説を厳密に検証するように設計されています。EDAで見つかった知見を、同じデータセット上で確定的な結論として扱ってしまうと、データドレッジング(p-hacking)や過学習を引き起こす原因になります。
import numpy as np
import pandas as pd
# 1. EDA Phase: Open-ended discovery and hypothesis generation
df = pd.DataFrame({'engagement_score': np.random.normal(50, 10, 1000)})
summary = df['engagement_score'].describe()
# 2. Confirmatory Phase: Testing pre-registered hypothesis on fresh test/experiment data
# (e.g., two-sample t-test with fixed significance level alpha = 0.05)
# Scenario: Predicting whether a user will cancel their subscription (is_churned)
# LEAKY FEATURE:
# 'cancellation_survey_submitted' -> Occurs after the decision to churn has executed.
# LEGITIMATE FEATURE:
# 'login_count_last_30_days' -> Observed strictly before the prediction cut-off date.
from sklearn.model_selection import train_test_split
from sklearn.datasets import make_classification
X, y = make_classification(n_samples=1000, n_features=10, random_state=42)
# Step 1: Split off final test set (20%)
X_dev, X_test, y_dev, y_test = train_test_split(
X, y, test_size=0.20, random_state=42, stratify=y
)
# Step 2: Split remaining development data into train (75% of dev = 60% total) and validation (25% of dev = 20% total)
X_train, X_val, y_train, y_val = train_test_split(
X_dev, y_dev, test_size=0.25, random_state=42, stratify=y_dev
)
print(f"Train size: {len(X_train)}, Val size: {len(X_val)}, Test size: {len(X_test)}")
8ノーススターメトリック(North Star Metric)とは何ですか?また、プロダクトの健全性を評価する際、効果的なノーススターと虚栄の指標(vanity metric)をどのように見分けますか?
ノーススターメトリック(NSM: North Star Metric)とは、持続可能なビジネス成果を推進しつつ、プロダクトが顧客に提供する中核的な価値を最も的確に捉えた主要指標です(例: Airbnbにおける「予約成立宿泊数」や、音楽プラットフォームにおける「週間アクティブ再生時間」)。これは、プロダクトチームの意識を表層的な成長ではなく長期的な顧客価値に一致させる役割を果たします。
効果的なノーススター指標と虚栄の指標を見分けるポイントは以下のとおりです。
1. 価値との整合性: 効果的なノーススター指標は実際のユーザーの利便性やアクティブなエンゲージメントを反映します。一方、虚栄の指標は、ユーザーが即座に離脱していても増加しうる表層的なボリューム(登録ユーザー総数、累計アプリダウンロード数、生のページビュー数など)を測定してしまいます。
2. アクション可能性と相関性: 効果的なノーススター指標は、ユーザーリテンション(継続率)、プロダクトの健全性、収益化と相関し、プロダクト品質の向上に直接反応します。虚栄の指標は真のリテンションやビジネスの健全性と結び付けられないことが多く、表層的な水増しが起こりやすい傾向があります。
Platform: Vacation Rental App
Vanity Metric: Cumulative app downloads (increases continuously even if 95% of users uninstall immediately).
North Star Metric: Nights booked per active user (reflects actual value exchange between guests and hosts).
# Data Drift (Covariate Shift):
# User demographics or devices change (P(X) shifts),
# but genuine vs fraudulent transaction patterns remain identical (P(Y|X) unchanged).
# Concept Drift:
# Fraudsters adapt tactics to mimic normal shopping patterns;
# identical feature inputs now have a higher probability of fraud (P(Y|X) shifts).
母数(母集団パラメータ)とは、母集団全体における固定的で通常は未知の数値的特性です(真の母平均 μ や母比率 p など)。標本統計量とは、未知の母数を推定するために観測された標本データから計算される数値の要約です(標本平均 x̄ や標本比率 p̂ など)。標本変動とは、同一母集団から抽出された異なる無作為標本間で標本統計量が自然に変動することを指します。標本変動が存在するため、単一の標本指標には常にランダムな誤差が含まれ、真の母数と完全に一致することは稀です。指標を解釈する際、この変動を考慮しないとノイズを真の変化と誤認することになります。分析者は、観測された差が本物であると結論付ける前に、標準誤差、信頼区間、または仮説検定を用いて推定の不確実性を定量化する必要があります。
import numpy as np
# True population parameter (mean = 50, standard deviation = 10)
pop_mean = 50.0
pop_std = 10.0
# Draw multiple independent samples of size n=30
np.random.seed(42)
sample_means = [np.mean(np.random.normal(pop_mean, pop_std, size=30)) for _ in range(5)]
for i, sm in enumerate(sample_means, 1):
print(f"Sample {i} Mean (Statistic): {sm:.2f} | Error: {sm - pop_mean:+.2f}")
WITH user_cohorts AS (
SELECT
user_id,
DATE_TRUNC('week', signup_time) AS cohort_week,
signup_time,
acquisition_channel
FROM users
-- Exclude cohorts that have not reached 7 full days of maturity
WHERE signup_time <= CURRENT_DATE - INTERVAL '7 days'
),
user_activity AS (
SELECT DISTINCT
c.user_id,
c.cohort_week,
c.acquisition_channel,
1 AS retained_d7
FROM user_cohorts c
JOIN events e ON c.user_id = e.user_id
WHERE e.event_time >= c.signup_time + INTERVAL '7 days'
AND e.event_time < c.signup_time + INTERVAL '8 days'
)
SELECT
c.cohort_week,
c.acquisition_channel,
COUNT(c.user_id) AS cohort_size,
COUNT(a.user_id) AS d7_active_users,
ROUND(100.0 * COUNT(a.user_id) / COUNT(c.user_id), 2) AS d7_retention_pct
FROM user_cohorts c
LEFT JOIN user_activity a ON c.user_id = a.user_id
GROUP BY 1, 2
ORDER BY 1 DESC, 2;